RT Journal Article SR Electronic T1 Socio-Demographic Biases in Medical Decision-Making by Large Language Models: A Large-Scale Multi-Model Analysis JF medRxiv FD Cold Spring Harbor Laboratory Press SP 2024.10.29.24316368 DO 10.1101/2024.10.29.24316368 A1 Omar, Mahmud A1 Soffer, Shelly A1 Agbareia, Reem A1 Bragazzi, Nicola Luigi A1 Apakama, Donald U. A1 Horowitz, Carol R A1 Charney, Alexander W A1 Freeman, Robert A1 Kummer, Benjamin A1 Glicksberg, Benjamin S A1 Nadkarni, Girish N A1 Klang, Eyal YR 2024 UL http://medrxiv.org/content/early/2024/10/30/2024.10.29.24316368.abstract AB Large language models (LLMs) are increasingly integrated into healthcare but concerns about potential socio-demographic biases persist. We aimed to assess biases in decision-making by evaluating LLMs’ responses to clinical scenarios across varied socio-demographic profiles. We utilized 500 emergency department vignettes, each representing the same clinical scenario with differing socio-demographic identifiers across 23 groups—including gender identity, race/ethnicity, socioeconomic status, and sexual orientation—and a control version without socio-demographic identifiers. We then used Nine LLMs (8 open source and 1 proprietary) to answer clinical questions regarding triage priority, further testing, treatment approach, and mental health assessment, resulting in 432,000 total responses. We performed statistical analyses to evaluate biases across socio-demographic groups, with results normalized and compared to control groups. We find that marginalized groups—including Black, unhoused, and LGBTQIA+ individuals—are more likely to receive recommendations for urgent care, invasive procedures, or mental health assessments compared to the control group (p < 0.05 for all comparisons). High-income patients were more often recommended advanced diagnostic tests such as CT scans or MRI, while low-income patients were more frequently advised to undergo no further testing. We observed significant biases across all models, both proprietary and open source regardless of the model’s size. The most pronounced biases emerged in mental health assessment recommendations. LLMs used in medical decision-making exhibit significant biases in clinical recommendations, perpetuating existing healthcare disparities. Neither model type nor size affects these biases. These findings underscore the need for careful evaluation, monitoring, and mitigation of biases in LLMs to ensure equitable patient care.Competing Interest StatementThe authors have declared no competing interest.Funding StatementThis study did not receive any fundingAuthor DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesI confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesAll data produced in the present study are available upon reasonable request to the authors