PT  - JOURNAL ARTICLE
AU  - Nakamura, Yuta
AU  - Kikuchi, Tomohiro
AU  - Yamagishi, Yosuke
AU  - Hanaoka, Shouhei
AU  - Nakao, Takahiro
AU  - Miki, Soichiro
AU  - Yoshikawa, Takeharu
AU  - Abe, Osamu
TI  - ChatGPT for automating lung cancer staging: feasibility study on open radiology report dataset
AID  - 10.1101/2023.12.11.23299107
DP  - 2023 Jan 01
TA  - medRxiv
PG  - 2023.12.11.23299107
4099  - http://medrxiv.org/content/early/2023/12/13/2023.12.11.23299107.short
4100  - http://medrxiv.org/content/early/2023/12/13/2023.12.11.23299107.full
AB  - Objectives CT imaging is essential in the initial staging of lung cancer. However, free-text radiology reports do not always directly mention clinical TNM stages. We explored the capability of OpenAI’s ChatGPT to automate lung cancer staging from CT radiology reports.Methods We used MedTxt-RR-JA, a public de-identified dataset of 135 CT radiology reports for lung cancer. Two board-certified radiologists assigned clinical TNM stage for each radiology report by consensus. We used a part of the dataset to empirically determine the optimal prompt to guide ChatGPT. Using the remaining part of the dataset, we (i) compared the performance of two ChatGPT models (GPT-3.5 Turbo and GPT-4), (ii) compared the performance when the TNM classification rule was or was not presented in the prompt, and (iii) performed subgroup analysis regarding the T category.Results The best accuracy scores were achieved by GPT-4 when it was presented with the TNM classification rule (52.2%, 78.9%, and 86.7% for the T, N, and M categories). Most ChatGPT’s errors stemmed from challenges with numerical reasoning and insufficiency in anatomical or lexical knowledge.Conclusions ChatGPT has the potential to become a valuable tool for automating lung cancer staging. It can be a good practice to use GPT-4 and incorporate the TNM classification rule into the prompt. Future improvement of ChatGPT would involve supporting numerical reasoning and complementing knowledge.Clinical relevance statement ChatGPT’s performance for automating cancer staging still has room for enhancement, but further improvement would be helpful for individual patient care and secondary information usage for research purposes.Key pointsChatGPT, especially GPT-4, has the potential to automatically assign clinical TNM stage of lung cancer based on CT radiology reports.It was beneficial to present the TNM classification rule to ChatGPT to improve the performance.ChatGPT would further benefit from supporting numerical reasoning or providing anatomical knowledge.Competing Interest StatementThe authors have declared no competing interest.Funding StatementThis study did not receive any funding.Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesI confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesAll data produced are available online at https://sociocom.naist.jp/medtxt/rr/ https://sociocom.naist.jp/medtxt/rr/ cTNMclinical TNM stageNLPnatural language processingAPIapplication programming interface