PT - JOURNAL ARTICLE AU - Bera, Kaustav AU - Gupta, Amit AU - Jiang, Sirui AU - Berlin, Sheila AU - Faraji, Navid AU - Tippareddy, Charit AU - Chiong, Ignacio AU - Jones, Robert AU - Nemer, Omar AU - Nayate, Ameya AU - Tirumani, Sree Harsha AU - Ramaiya, Nikhil TI - Assessing Performance of Multimodal ChatGPT-4 on an image based Radiology Board-style Examination: An exploratory study AID - 10.1101/2024.01.12.24301222 DP - 2024 Jan 01 TA - medRxiv PG - 2024.01.12.24301222 4099 - http://medrxiv.org/content/early/2024/01/13/2024.01.12.24301222.short 4100 - http://medrxiv.org/content/early/2024/01/13/2024.01.12.24301222.full AB - Objective To evaluate the performance of multimodal ChatGPT 4 on a radiology board-style examination containing text and radiologic images.sMaterials and Methods In this prospective exploratory study from October 30 to December 10, 2023, 110 multiple-choice questions containing images designed to match the style and content of radiology board examination like the American Board of Radiology Core or Canadian Board of Radiology examination were prompted to multimodal ChatGPT 4. Questions were further sub stratified according to lower-order (recall, understanding) and higher-order (analyze, synthesize), domains (according to radiology subspecialty), imaging modalities and difficulty (rated by both radiologists and radiologists-in-training). ChatGPT performance was assessed overall as well as in subcategories using Fisher’s exact test with multiple comparisons. Confidence in answering questions was assessed using a Likert scale (1-5) by consensus between a radiologist and radiologist-in-training. Reproducibility was assessed by comparing two different runs using two different accounts.Results ChatGPT 4 answered 55% (61/110) of image-rich questions correctly. While there was no significant difference in performance amongst the various sub-groups on exploratory analysis, performance was better on lower-order [61% (25/41)] when compared to higher-order [52% (36/69)] [P=.46]. Among clinical domains, performance was best on cardiovascular imaging [80% (8/10)], and worst on thoracic imaging [30% [3/10)]. Confidence in answering questions was confident/highly confident [89%(98/110)], even when incorrect There was poor reproducibility between two runs, with the answers being different in 14% (15/110) questions.Conclusion Despite no radiology specific pre-training, multimodal capabilities of ChatGPT appear promising on questions containing images. However, the lack of reproducibility among two runs, even with the same questions poses challenges of reliability.Competing Interest StatementThe authors have declared no competing interest.Funding StatementThe study did not receive any funding.Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesI confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesAll data produced in the present work are contained in the manuscript