SHORT COMMUNICATION
Jonathan SHAPIRO1, Anna LYAKHOVITSKY2,3, Tamar FREUD4, Felix PAVLOTSKY2,3, Ziad KHAMAYSI5,6, Yulia VALDMAN-GRINSHPOUN7, Roni DODIUK-GAD5,8,9, Ilan GOLDBERG3,10, Arieh INGBER11, Baruch KAPLAN12 and Emily AVITAN-HERSH5,6
1Maccabi Healthcare Services, Tel Aviv-Yafo, Israel, 2Department of Dermatology, Sheba Medical Center, Tel Hashomer, Ramat Gan, Israel, 3School of Medicine, Faculty of Medical and Health Sciences, Tel Aviv University, Tel Aviv, Israel, 4Ben-Gurion University of the Negev, Beer-Sheva, Israel, 5Department of Dermatology, Rambam Health Care Campus, Haifa, Israel, 6Technion Faculty of Medicine, Haifa, Israel, 7Soroka Medical University Center, Beer-Sheva, Israel, 8Department of Dermatology, Emek Medical Center, Afula, Israel, 9Division of Dermatology, Department of Medicine, University of Toronto, Canada, 10Division of Dermatology, Tel-Aviv Sourasky Medical Center, Tel-Aviv, Israel, 11Hadassah Medical Center, Jerusalem, Israel, and 12Adelson School of Medicine, Ariel University, Ariel, Israel. E-mail: jonmidi@gmail.com
Citation: Acta Derm Venereol 2025; 105: adv41208. DOI: https://doi.org/10.2340/actadv.v105.41208.
Copyright: © 2025 The Author(s). Published by MJS Publishing, on behalf of the Society for Publication of Acta Dermato-Venereologica. This is an Open Access article distributed under the terms of the Creative Commons Attribution-NonCommercial 4.0 International License (https://creativecommons.org/licenses/by-nc/4.0/).
Submitted: Jul 20, 2024; Accepted after revision: Oct 28, 2024; Published: Jan 3, 2025
Competing interests and funding: The authors have no conflicts of interest to declare.
The Chat Generative Pre-Trained Transformer (ChatGPT) series is pivotal in natural language and image processing (1). ChatGPT has shown near-passing results in medical licensing exams, including dermatology (2–4). An assessment of ChatGPT-3.5 for the American Board of Dermatology Applied Exam found 40% of its questions accurate and suitably complex (5). ChatGPT-4 advances further with improved linguistic processing, deeper subject understanding, and a broader knowledge base, potentially improving its question-generation capability.
The Israeli Dermatology Board exam preparation involves a multi-stage process. Based on the textbook “Dermatology, 4th Edition”, by Bolognia et al. (6), committee members create 150 multiple-choice questions, each based on a different chapter. The chair reviews these questions for accuracy and structure. The committee then discusses each question and stratifies questions by difficulty. Key rules include having one correct answer, avoiding “all of the above“, “none of the above”, and double negatives, and ensuring answers are syllabus-based. The exam also features complex clinical cases requiring diagnoses based on descriptions and images, and the questions relate to different clinical or laboratory characteristics of the diagnosis.
This study assesses the effectiveness of ChatGPT-4 in producing accurate and contextually relevant examination content for dermatology board exams.
Twelve thematic areas were randomly chosen from the textbook “Dermatology, 4th Edition”, by Jean L. Bolognia, Julie V. Schaffer, and Lorenzo Cerroni (6). The text of each specific chapter was copied into a Word document and securely imported into the paid version of ChatGPT-4, which was commercially available between 27 December 2023, and 3 January 2024. The “Chat & History Training” parameter in ChatGPT-4’s data control settings was disabled to prevent the data from being used for training or stored on its servers. Chats were automatically deleted upon completion, with no option for recovery. Subsequently, the model was tasked to generate multiple-choice questions. The prompt was refined after a systematic process of trial and error and is detailed in Appendix S1. The following final prompt version was consistently used for all the subjects: “Based only on the Word document I uploaded, ask extremely hard complicated, and very diverse questions including regular and clinical questions and a two-step thought process and provide the answer after every question and write at what page in the document I uploaded I can find the answer. If the question requires a two-step thought process where the physician must first deduce the diagnosis from the clinical presentation before answering the specific question, don’t mention the diagnosis in the questions and add the diagnosis to the answer in a separate line. The questions should be multiple choice numbered questions.”. The prompt and the questions were both in English.
Eight board-certified dermatology experts reviewed the questions. Of those, 5 (FP, YVG, IG, AI, and EAH) are long-term members of the Israeli board exam committee (10, 4, 8, 8, and 7 years, respectively). Two authors chaired the committee (FP, AI) and 1 is the current chair (EAH).
Each questionnaire was assessed by 2 reviewers, of which at least 1 was a long-term committee member. All questions were evaluated as “Suitable“, which were further graded by difficulty, or “Not Suitable“, which were categorized based on the reason. In cases of disagreement, mutual consultations were aimed at reconciling differences in scoring. Reviewers also recorded the time spent reviewing each exam and estimated how long it would have taken to write the same number of appropriate questions.
Statistical analysis was primarily descriptive. Categorical variables were presented as frequency and percentage. Inter-rater reliability was calculated utilizing Cohen’s Kappa. All analyses were performed with IBM SPSS statistic software version 29.0 (IBM Corp, Armonk, NY, USA). P<0.05 was chosen as the significance level.
ChatGPT-4 generated 402 questions, with 208 (51.7%) deemed acceptable by at least 1 reviewer. However, only 72 questions (18%) were accepted by both reviewers. After consensus discussions, 53 of the 136 initially disputed questions were approved, resulting in a total of 125 questions deemed suitable for the exam. The suitable questions were classified as 51 (40.8%) easy, 45 (36%) medium-difficulty, and 29 (23.2%) hard. The main issues with unsuitable questions included questions that contained errors or improperly structured or with potential for an appeal (118 questions, 27.8%) and excessive simplicity (113 questions, 28.1%).
Table I provides a breakdown of the generated questions by subject area. Biopsy techniques and B-cell lymphoma had the highest rates of suitable questions (63–65%). In addition, 37 questions were 2-stage complicated questions. Of those 7 were determined as appropriate (18.9%).
In 19 of the 24 assessments, the reviewers acknowledged that using ChatGPT-4 could potentially reduce the time needed by up to 55 min per question (range –110 to –55). Table II presents the time invested to review each subject and the estimated duration for designing suitable questions. Most reviewers rated the platform as useful and exhibited their willingness to employ it in the future.
In our cohort, the inter-rater reliability was low, indicating a generally low level of agreement before consensus (Table III). The Kappa values were highest in the vasculitis and HPV chapters. Of the 136 disputes, 55 (40.4%) arose from 1 reviewer finding the question too easy, while 48 (35.3%) involved errors or poor structure.
ChatGPT has gained significant popularity for its natural language processing and content generation capabilities. Given the complexity of structuring board exams, which demands consistency, proper question structure, and a balanced mix of difficulty levels and clinical scenarios, we explored ChatGPT-4’s ability to generate suitable multiple-choice questions for dermatology board exams. This study extends previous work with ChatGPT-3.5 (5, 7–10) by increasing the number of questions and incorporating two-step reasoning tasks, such as diagnostic deductions and follow-up actions (e.g., “What would be your next step?’), to evaluate the model’s performance comprehensively.
In generating multiple-choice questions, initial attempts yielded overly simple questions. Therefore, we revised the prompt to request highly complex questions with two-step reasoning. Despite this, over a third of the questions were still deemed too easy. As not all easy questions are inappropriate, we included 51 such questions, recognizing the difficulty in distinguishing “too easy” from “easy but acceptable”. Of the accepted questions, 29 were classified as hard, and 7 involved two-stage reasoning. This underscores the challenge of using ChatGPT-4 to produce suitably complex questions that accurately assess clinical scenarios. Additionally, 118 questions were flagged due to ambiguous wording or multiple correct answers. This highlights a key challenge for ChatGPT-4: ensuring clarity and a single correct answer to prevent disputes. Effective examination design requires questions to be not only factually accurate but also clear and precise, a standard that remains challenging for AI platforms to meet.
We assessed ChatGPT-4’s ability to produce diverse questions and found that 2.2% of the questions were repeated, indicating limited novelty compared with human-generated questions. Students and residents might also generate questions that will be similar to those appearing on exams, highlighting a potential issue with using the platform. However, the reviewers recognized the educational value of assessing AI-generated questions, noting that this process fosters deeper engagement with the curriculum, which may be beneficial for medical students. This suggests that, with further refinement, AI could be adapted for various levels of medical education, enhancing learning outcomes across different stages (7, 10).
Future AI iterations in question generation should enhance algorithms to better assess question complexity and reduce ambiguities that may lead to appeals. This requires integrating expert feedback into the AI training process to align outputs with board-certified professionals’ expectations. Collaboration between AI developers and educational experts is essential for advancing AI capabilities, ensuring outputs meet educational standards and learning objectives, and potentially improving both question-generation efficiency and educational support in medical training.
In conclusion, this analysis highlights both the potential and limitations of ChatGPT-4 in generating questions of varying difficulty and complex clinical scenarios. Key constraints include a significant proportion of overly simplistic questions and inaccurate distractor options. With improved training and contextual understanding, AI tools could better leverage their potential, addressing current limitations and generating diverse questions across subjects. Thus, ChatGPT-4 currently emerges as a supplementary tool for dermatology board exam preparation and may become more effective with forthcoming modifications.