Stephan Rau, Alexander Rau, Johanna Nattenmßller, Anna Maria Fink ¡ 7 authors
BACKGROUND: We investigated the potential of an imaging-aware GPT-4-based chatbot in providing diagnoses based on imaging descriptions of abdominal pathologies. METHODS: Utilizing zero-shot learning via the LlamaIndex framework, GPT-4 was enhanced using the 96 documents from the Radiographics Top 10 Reading List on gastrointestinal imaging, creating a gastrointestinal imaging-aware chatbot (GIA-CB). To assess its diagnostic capability, 50 cases on a variety of abdominal pathologies were created, comprising radiological findings in fluoroscopy, MRI, and CT. We compared the GIA-CB to the generic GPT-4 chatbot (g-CB) in providing the primary and 2 additional differential diagnoses, using interpretations from senior-level radiologists as ground truth. The trustworthiness of the GIA-CB was evaluated by investigating the source documents as provided by the knowledge-retrieval mechanism. Mann-Whitney U test was employed. RESULTS: The GIA-CB demonstrated a high capability to identify the most appropriate differential diagnosis in 39/50 cases (78%), significantly surpassing the g-CB in 27/50 cases (54%) (p = 0.006). Notably, the GIA-CB offered the primary differential in the top 3 differential diagnoses in 45/50 cases (90%) versus g-CB with 37/50 cases (74%) (p = 0.022) and always with appropriate explanations. The median response time was 29.8 s for GIA-CB and 15.7 s for g-CB, and the mean cost per case was $0.15 and $0.02, respectively. CONCLUSIONS: The GIA-CB not only provided an accurate diagnosis for gastrointestinal pathologies, but also direct access to source documents, providing insight into the decision-making process, a step towards trustworthy and explainable AI. Integrating context-specific data into AI models can support evidence-based clinical decision-making. RELEVANCE STATEMENT: A context-aware GPT-4 chatbot demonstrates high accuracy in providing differential diagnoses based on imaging descriptions, surpassing the generic GPT-4. It provided formulated rationale and source excerpts supporting the diagnoses, thus enhancing trustworthy decision-support. KEY POINTS: ⢠Knowledge retrieval enhances differential diagnoses in a gastrointestinal imaging-aware chatbot (GIA-CB). ⢠GIA-CB outperformed the generic counterpart, providing formulated rationale and source excerpts. ⢠GIA-CB has the potential to pave the way for AI-assisted decision support systems.
Open access
Artificial Intelligence in Healthcare and Education
Figure: emergency medicine, ACEP, AAEM, emergency physicians, American College of Emergency Physicians, American Academy of Emergency Medicine, CMGs, demographics, private equity, consolidationFigureEmergency physicians are strong-willed, smart, distractible, and accustomed to being in charge, but we're a bit like herding catsâchallenging to unite. Nonetheless, our professional organizations are tasked with solving the problems of our specialty by uniting physicians with common goals. Understanding why physicians join or leave these organizations, specifically, why the physicians on the cusp are thinking of changing their membership status, is crucial to broadening organizational appeal and identifying unifying concerns. Physicians with strong opinionsâpositive or negativeâabout specific professional organizations are not likely to be on the cusp of changing membership status. Members and non-members on the margins, however, are potentially high-yield targets for gaining or retaining members. Recognizing the concerns of this group may be key to a professional organization's success. So, I decided to look at the American College of Emergency Physicians and the American Academy of Emergency Medicine. Importance of the Margin âThis is exactly what the net promoter score is,â said Mollie Pillman, ACEP's senior vice president of membership engagement. âThe theory behind it is that you have your promoters, your detractors, and your passives in the middle, and you really want to go after the passives because they're the ones whose minds you'll be able to change.â ACEP's Executive Director Sue Sedory agreed, noting that the margin is key for the specialty to stay strong. âIf we can get more people who just want to engage and have a stake in the future of the specialty and then help them one by one, which is what our commitment has been, [that] is really key.â AAEM's president, Jonathan Jones, MD, said it's easy to attract people who already agree with you, but that he didn't want members just for members' sake. âWhat I do want is people that believe in our basic principles but have a different perspective, that have a different background, that have a different plan.â ACEP and AAEM said they recognize the need for a more granular understanding of marginal members and non-members, but neither engages in mathematical modeling of this problem. âWe are about nine to 10 months into association analytics implementation, which is a platform built for associations that pulls in all of your data from events and from finance and from the membership database, and it consolidates all of it together,â Ms. Sedory said. âSo hopefully we'll be able to answer that better in the future.â AAEM has yet to invest in such a system either, though it is something the academy has considered. Both organizations said they also recognize the importance of incorporating a diversity of opinions. âACEP is taking proactive steps to integrate member and non-member feedback into its strategic planning process, evaluation of new offerings, and communication effectiveness,â the college said in a statement. âWe regularly conduct focus groups which are carefully chosen to represent a wide range of demographics and perspectives, and test new ideas and benefits with diverse groups before they are made available to a wider audience.â Dr. Jones said AAEM strives not to become stuck in its own bubble by reaching out to physicians in academia, CMGs, and democratic groups. âIf we have people from different practice environments, that's the best way that we can try not to become an echo chamber,â he said. âAnd obviously we want a diverse membership, not just from practice environments but geographically [and by] gender and race.â Proof of Concept I recently surveyed emergency physicians on Facebook about their thoughts on professional organizations. The total response number was low, just more than 100, which is not nearly high enough to generalize about emergency physician preferences. Selection bias was also a real concern with a smaller group acquired this way. My intent with these data was not generalization, however, but to demonstrate how a professional organization might assess their marginal members and non-members. If a group this size can be effectively segmented with a relatively small amount of data, then the same thing can be done more easily and effectively with larger sample sizes and larger amounts of data. I segmented the group based on a specific question: How much do you value membership in your professional organizations? I used k-medoids clustering (an unsupervised machine learning process that breaks data into k clusters by iteratively picking samples within data to serve as the cluster's best median, making it more robust to outliers than the similar k-means clustering process) to sort respondents into detractors (respondents who view the organization very poorly), promoters (respondents who view the organization very positively), and passives. I asked about the issues that professional organizations should prioritize more to see if there were differences of opinions based on the perceived value cluster allocation. These topics were selected based on feedback from interviewees and ACEP's latest annual report. Respondents had specific concerns (some different, some the same) about ACEP and AAEM that are more pressing for passives or marginal members than for other clusters. I ran a two-way ordinal analysis of variance testing of ACEP and AAEM clustering to confirm statistical significance on the (visible) differences. Marginal members of both organizations considered corporate consolidation and private equity in medicine to be an issue deserving of more attention.FigureI then flipped the tables and used only the respondents' age and their responses to the priority questions to determine how far apart respondents were using a metric called the Gower distance (to account for the ordinal data of the Likert scale). I chose to use three clusters (based on maximization of a metric called the silhouette width; this assesses the uniqueness of the clusters), and used a similar process to k-medoids clustering to form groups of respondents based only on what priorities they considered important, this time without knowledge of membership status or value. Clustering was able to sort respondents into groups that seemed to favor AAEM membership (cluster 1), favor ACEP membership (cluster 2), or eschew both organizations (cluster 3). Average perceived value of the organizations also followed this characterization. Cluster 1 valued AAEM membership the most and ACEP membership second, and the opposite was true for cluster 2. A United Specialty Both ACEP and AAEM report caring a great deal about members who may be on the margins; there are ways to find these members and determine what matters most to them. If organizations want to avoid becoming entrenched echo chambers, they must actively seek out marginal members and non-members and attempt to incorporate them in meaningful ways. A united specialty is clearly important, but marginal members do not want the same things as their nonmarginal counterparts; it is the responsibility of our professional societies to achieve real unity in critical areas by meaningfully incorporating the differences of opinion of others. There is likely a subset of physicians who is generally disenchanted, but ACEP's marginal population appears to be AAEM members and vice versa. This paints a hopeful picture in my mind: Reaching out and incorporating those with different ideas may be easier than we think. Even for emergency physicians. DR. BELANGER is the chief data officer of TotalCare (totalcare.us), the chair-elect of the American College of Emergency Physicians Workforce Section, and an emergency physician in McKinney, TX.
Dr. Welch: is a fellow with Intermountain Institute for Health Care Delivery Research, an emergency physician with Utah Emergency Physicians, and a member of the board of the Emergency Department Benchmarking Alliance. Dr. Taylor is a physician executive for Microsoft Corporation's Health Solutions Group and the principle promoter of the national Emergency Department of the Future project. Dr. Cheung is a Malcolm Baldrige National Quality Award Examiner, former faculty of the Johns Hopkins Center for Innovation in Quality Patient Care and the Quality and Safety Research Group, and a member of ACEP's Quality and Performance Committee.Should board certified emergency physicians be required to earn certain âmerit badgeâ credentials for hospital privileges? Should such certifications be required for activities like EMS base station medical control or specialty center designation? Hospitals, state EMS authorities, and specialty center designation bodies are increasingly turning to âmerit badgesâ (see table) as proxies for competency in various specialty areas within emergency medicine. Does this make any sense? And where will this end? Will emergency physicians ultimately be required to hold certificates or earn CME hours in geriatric medicine, psychiatry, ethics, and ingrown toenail removal? An Internet search for âmerit badge medicineâ results in web sites for the Boy Scouts. Has it come to this? Should we wear merit badge sashes during our shifts? It's time to insert a bit of common sense into this madness. In the early days of the specialty, there were few training programs, and the first certifying examination in emergency medicine did not occur until 1980. As a result, most early emergency physicians gained on-the-job experience, and those who became board certified did so through the practice pathway. Since that time, training programs have expanded dramatically, and the only recognized path to board certification is now through emergency medicine residency training and the American Board of Medical Specialties' certifying bodies. Ironically, despite 30 years of formal emergency medicine board certification, âmerit badgeâ requirements have re-emerged as a phenomenon. Are these efforts unwarranted, unnecessary, or even counterproductive? And what is driving this trend? The American College of Emergency Medicine and the American Academy of Emergency Medicine are clear on the use of âmerit badges,â and these positions can be used as support exemption from merit badge requirements: The AAEM membership card notes that a fellow is board certified in emergency medicine and therefore has advanced resuscitation expertise in pediatric, trauma and cardiac care. The 1999 ACEP policy, âUse of Short Courses in Emergency Medicine as Criteria for Privileging or Employment,â strongly discourages the use of certificates in subareas of emergency medicine as requirements for privileges or employment. (http://bit.ly/ShortCourse.) Nevertheless, the growing list of merit badges required for hospital credentialing and privileging has become cumbersome. Resources for continuing medical education for physicians can be a zero-sum game. The time and cost for merit badge achievement will inevitably result in other, perhaps more important, educational areas being ignored. Most states and specialty societies require CME for maintenance of licensure and membership. On average, this requirement is 20 to 50 credit hours a year, and can often require up to two weeks away from work (worth $10,000 to $20,000) and $5000 to $8,000 for tuition and travel expenses. While making a financial argument against merit badges can be risky (i.e., âdoctors make lots of moneyâ), it can be done successfully. Arizona ACEP avoided state-mandated merit badges by pointing out the total annual cost for particular requirements. An effort to require PALS by all ED staff (including nurses) was defeated when it was suggested it should be funded as a state-mandated program at an annual cost of more than $8 million. This idea quietly went away. As a relatively young specialty, emergency medicine may suffer from a lack of understanding by the community, hospital administrators, and regulators on the breadth of expertise and training of board certified emergency physicians. Even among emergency physicians, considerable confusion exists over the Emergency Medicine Continuous Certification program requirements. These requirements, built on residency training and certification, are perhaps the best argument against merit badges. With the significant investment in initial training and the ongoing and increasing Maintenance of Certification requirements, merit badges are simply unnecessary. There is no evidence that any of these individual credentials correlate with improved clinical care. In contrast, one study demonstrated a quality correlation between board certification and quality measures. (Arch Intern Med 2010;170[16]:1442.) In addition, the âcredentials equal qualityâ mental model is founded on a false premise, that putting more knowledge in the head of the practitioner will improve care. In the Agency for Healthcare Research and Quality whitepaper, âMistake Proofing the Design of Health Care Processesâ (www.ahrq.gov), the authors note that one of the biggest failed mental models in health care has been to assume that if we could put more into the practitioner's memory, we could avoid mistakes and provide better care. The premise is faulty because the human memory is fallible. The passage of time erases these efforts from memory if not used regularly. An example is pediatric resuscitation credentialing where pediatric critical care occurs only once every 30,000 to 40,000 ED visits. So the average PALS certified practitioner in a community hospital will go years or even decades before using the information, which is unlikely to be retrievable when actually needed. No other specialty requires physicians to jump through as many hoops to put on a white coat and practice medicine. Five arguments can be used against merit badge requirements for board certified emergency physicians: ACEP and AAEM have strong policy statements against merit badge requirements. Maintaining these credentials is growing more burdensome, expensive, and crowding out more important CME activities. The new ABMS Maintenance of Certification requirements supersede the need for merit badges. There is no proven quality correlation with merit badge requirements. The mental model of âknowledge in the headâ is a false premise. Comments about this article? Write to EMN at[email protected]. Click and Connect!Access the links in this article by reading it onwww.EM-News.com. Typical Merit Badges Basic Life Support (BLS) Advanced Cardiac Life Support (ACLS) Advanced Trauma Life Support (ATLS) Advanced Pediatric Life Support (APLS) Pediatric Advanced Life Support (PALS) Advanced Airway Courses Procedural Sedation Courses Ultrasound Training and Credentialing Hazardous Materials Training (HAZMAT) The Joint Commission's Ongoing Practice Performance Evaluation (OPPE) 16 hours a year of trauma CME (for physicians practicing in American College of Surgeon Certified Trauma Centers) Various state-specific requirements (pain management, geriatric medicine, end-of-life care, infectious disease, risk management, sexual assault, domestic violence, cultural competence, appropriate prescribing, medical jurisprudence, ethics in medicine) CME Requirements Read a state-by-state listing of CME requirements at http://bit.ly/CMErequirements.
This journal has recently adjusted its requirements for research papers, referring to the World Medical Associationâs Helsinki Declaration on Ethical Principles for Medical Research Involving Human Subjects.1,2 Many advances in medicine have been attained by research involving patients. Individual patientsâ best interests have sometimes been sacrificed for the benefit of future patientsâ health. Most countries have established legal regulations and procedures, based upon this declaration, along with institutional review boards (IRBs). Such universal requirements do not exist for education research or social science research in general. For this reason, many countries, but not all, have included medical education research in their ethical review procedures, which are designed for medical research. Human subjects here are usually students or residents, but teachers and patients can also be involved. This journal will now require proof of ethical research conduct before publication. This step can be considered to represent an advance in research standards and may be taken up by other journals in the field. Like patients, learners are potentially vulnerable subjects of research, specifically if the investigators simultaneously exercise power as teachers or examiners. The Netherlands is among those countries that do not require ethical approval for medical education research. Dutch IRBs typically respond to submitted requests for review of education projects with statements like âexempt from ethical reviewâ because, firstly, patients are not involved and, secondly, no medical interventions are applied. So far, journals have accepted and published such statements without further question. Researchers often find this a comfortable stance as it avoids the bureaucratic burden of approval that has seriously hampered research elsewhere.3â5 For instance, I saw the recent 3-month stay in the UK of one of my research staff end without the planned interview and questionnaire project carried out, only because of a late and negative response from the IRB. Not that the project was unethical; the application simply lacked the requisite paperwork, despite the fact that extensive written and oral information was supplied. The question here is not whether ethical review in medical education research is justified â of course it is â but whether existing IRB procedures are most suitable for medical education research.6 Dutch researchers find âexempt from ethical approvalâ a comfortable stance as it avoids bureaucratic burden Given the requirements journals will put upon submissions, in terms of providing other proof of ethical conduct of medical education research if institutional review is not possible,1,7 independent ethical review will also become important in those countries without relevant procedures, such as the Netherlands. In a recent survey about experiences with IRBs among clinical course directors in the USA, it was suggested that national guidelines for ethics review would enhance transparency and stimulate inter-institutional research collaboration.8 When suggesting that the Netherlands Association for Medical Education take the lead in devising such guidelines, I began to wonder what ethical review of education research should look like, vis Ă vis the Helsinki Declaration. I concluded that education differs from health care in a number of aspects that might affect how ethics review should take place. Patients usually need care because of ill health. They are often very dependent on doctors and hospitals for their essential wellbeing. Students are also dependent as they must abide by the regulations of an institution and its teachers to pass necessary examinations, but they themselves are responsible for whether or not they enrol in particular courses and for whether they attain their self-chosen goals in life. They are, by far, not as dependent as patients. Next, students determine to a large extent the outcome of educational interventions, much more so than patients can determine the outcome of medical interventions. Personal study effort is a major determinant of academic success; medical treatment successes are determined to a far greater extent by health care providers. Students determine the outcome of educational interventions far more than patients can determine the outcome of medical interventions Medical research may involve risks of harm to a patientâs health, which, in some cases, may be serious and irreversible. This type of research requires the utmost caution. Harm in education research may be defined as the risk that less than optimal education is provided, resulting in less acquiral of knowledge and skills, and harm to academic progress. Harm to progress can be serious. I have witnessed a financial claim by a medical student who argued that an examiner had caused loss of income as a result of failing this student on tests and thus extending the required course length. The claim was not sustained and in this case no research was involved, but an educational experiment could be envisioned to carry a risk for such potential harm. In other cases, harm may imply psychological stress and discomfort. Still, this type of harm does not compare with potential physical harm resulting from medical experiments. Potential harm resulting from educational experiments does not compare with potential harm from medical experiments Furthermore, there is no clear distinction between administration and research purposes in data collection on student progress and programme quality. Schools collect data on student progress and the quality of educational processes as part of their core business. This type of data collecting is usually not subject to ethical approval, as registering test results and, to a lesser extent, registering studentsâ opinions of education are unavoidable. Enrolment in education must represent the subjectâs tacit approval that his or her personal data are registered. These data serve both students and educational quality. Reports based on aggregated data of student progress and educational quality can be considered to serve a necessary research aim; indeed, not using these data to improve education could be considered unethical. The necessity of gaining ethical approval or informed consent to use these data as part of research results for publication in a journal is questionable if harm to students is clearly not at stake. Researchers now sometimes label investigations as âevaluationâ in order to avoid the burden and delay incurred by a review procedure, but this does not seem a proper way to go. Rather, clear specifications of the types of data collection and uses that require ethical approval should guide researchers in their ethical conduct. Education is a process that cannot easily be stopped. Programmes must be offered to enrolled students and schools should continue to aim to deliver high-quality curricula. Many medical schools evaluate and try to improve their curricula, either gradually or by instigating large innovations all at once. Viewed from a research perspective, an educational method can be considered an intervention, just as a medical treatment is an intervention. At variance with medicine, medical education is only at a very early stage in the development of evidence-based practice. Education research â still â offers only little evidence to support a claim that one method is superior to another, despite the growing medical education literature.9,10 One reason for this is that students themselves determine to a large extent the effect of education.11 The counterside of the coin is that institutions can change their educational methods if they feel the need to do this, with limited risk for damage to the curriculum or harm to students. Letâs now compare curriculum development with education research. Consider a complete overhaul of the medical curriculum for a new student cohort, say, from a traditional to a problem-based learning (PBL) format. Does this âexperimentâ require ethical approval? Not likely. Not even when national examination results are compared with those of the previous cohort and published in support of the new development. This holds true if the results of students from two medical schools with such different curricula are compared. Next, consider a complete overhaul of the second year of a curriculum for only part of a student cohort. Students could be randomly assigned to either a PBL or a traditional programme, and their scores on national examinations could be compared. Would this âexperimentâ require ethical approval? Most likely it would in most countries. But what is the difference in terms of potential harm to students? Probably none. In general, schools and curricula differ in their educational methods and it is often hard to claim that students are better off with one school or method than another. Twenty years of external review of the eight medical schools in the Netherlands have not raised any serious conclusions that one school is clearly superior to another. Of course, checklists were used and the scores calculated showed differences on points, but no sound basis has ever emerged in support of real qualitative differences in outcome. This example shows how close educational development is to education research. The difference may only be that one instance is considered research, as it has a systematic project description, labels a new method as an intervention, formulates outcome measures and analyses data. What actually happens with students â the basis for ethical review â may be the same in the other instance. This PBL example may not sound very experimental because the format of PBL has been researched extensively, but such research has usually happened only after PBL has been introduced into a curriculum, not before. In other words, is it logical to require ethical approval when methods are compared within cohorts, but not to do so with scientifically less sound approaches, such as historical or inter-institutional comparisons? Providing a completely new educational method to a new cohort without any comparison does not require ethical approval, but it is day-to-day practice in many schools. Why require ethical approval when methods are compared within cohorts, but not historically or inter-institutionally? One important element of ethical research conduct concerns the obligation to provide subjects with the option not to be involved in the investigation, so that potential candidates are asked to give their informed consent to participation and provided with options for withdrawal. In education research in a field setting, participation in research often equals participation in education. For example, introducing a new knowledge test â in fact, most knowledge tests in educational settings can be considered as ânewâ instruments if test items have not been used before â cannot include an option for non-participation when a pass/fail decision must be based on such a test. Informed consent can be sought, but often there is no alternative available. In conclusion, education research should be carried out ethically and the journals that publish such studies have a responsibility to stimulate ethical research conduct. This includes the proper protection of the interests of any human subjects. However, the criteria with which we may evaluate the ethical conduct of education research are not necessarily equivalent to those required in medical research involving patients. Ethical research requires the maintenance of a reasonable balance, between yin and yang, so to speak, or between the risks of harm to subjects and the expected scientific yield of the investigation. This balance may well differ between medical and education research.12 I believe it would help to formulate specific criteria with which we can evaluate the ethics of medical education research. Pugsley and Dornan cite a helpful list of 10 ethical questions for research involving students, formulated by Cardiff University.7 This might represent a starting point for the development of such criteria and could result in a review procedure that both upholds ethical research standards and stimulates, rather than discourages, teachers and students to engage in medical education research.