Navigating Gynecological Oncology with Different Versions of ChatGPT: A Transformative Breakthrough or the Next Black Box Challenge?
| dc.authorid | 0000-0003-1019-3769 | |
| dc.authorid | 0000-0002-7234-3876 | |
| dc.contributor.author | Gungor, Nur Dokuzeylul | |
| dc.contributor.author | Esen, Fatih Sinan | |
| dc.contributor.author | Tasci, Tolga | |
| dc.contributor.author | Gungor, Kagan | |
| dc.contributor.author | Cil, Kaan | |
| dc.date.accessioned | 2025-05-10T19:41:09Z | |
| dc.date.issued | 2024 | |
| dc.department | İstanbul Medeniyet Üniversitesi | |
| dc.description.abstract | Introduction: The study evaluates the performance of large language model versions of ChatGPT - ChatGPT-3.5, ChatGPT-4, and ChatGPT-Omni - in addressing inquiries related to the diagnosis and treatment of gynecological cancers, including ovarian, endometrial, and cervical cancers. Methods: A total of 804 questions were equally distributed across four categories: true/false, multiple-choice, open-ended, and case-scenario, with each question type representing varying levels of complexity. Performance was assessed using a six-point Likert scale, focusing on accuracy, completeness, and alignment with established clinical guidelines. Results: For true/false queries, ChatGPT-Omni achieved accuracy rates of 100% for easy, 98% for medium, and 97% for complicated questions, higher than ChatGPT-4 (94%, 90%, 85%) and ChatGPT-3.5 (90%, 85%, 80%) (p = 0.041, 0.023, 0.014, respectively). In multiple-choice, ChatGPT-Omni maintained superior accuracy with 100% for easy, 98% for medium, and 93% for complicated queries, compared to ChatGPT-4 (92%, 88%, 80%) and ChatGPT-3.5 (85%, 80%, 70%) (p = 0.035, 0.028, 0.011). For open-ended questions, ChatGPT-Omni had mean Likert scores of 5.8 for easy, 5.5 for medium, and 5.2 for complex levels, outperforming ChatGPT-4 (5.4, 5.0, 4.5) and ChatGPT-3.5 (5.0, 4.5, 4.0) (p = 0.037, 0.026, 0.015). Similar trends were observed in case-scenario questions, where ChatGPT-Omni achieved scores of 5.6, 5.3, and 4.9 for easy, medium, and hard levels, respectively (p = 0.017, 0.008, 0.012). Conclusions: ChatGPT-Omni exhibited superior performance in responding to clinical queries related to gynecological cancers, underscoring its potential utility as a decision support tool and an educational resource in clinical practice. | |
| dc.identifier.doi | 10.1159/000543173 | |
| dc.identifier.issn | 2296-5270 | |
| dc.identifier.issn | 2296-5262 | |
| dc.identifier.pmid | 39689699 | |
| dc.identifier.scopusquality | Q3 | |
| dc.identifier.uri | https://doi.org/10.1159/000543173 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14730/10235 | |
| dc.identifier.wos | WOS:001404655900001 | |
| dc.identifier.wosquality | Q3 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | PubMed | |
| dc.language.iso | en | |
| dc.publisher | Karger | |
| dc.relation.ispartof | Oncology Research and Treatment | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/closedAccess | |
| dc.snmz | KA_WOS_20250302 | |
| dc.subject | ChatGPT | |
| dc.subject | Gynecological oncology | |
| dc.subject | Artificial intelligence | |
| dc.subject | Artificial intelligence accuracy | |
| dc.subject | Large language model | |
| dc.title | Navigating Gynecological Oncology with Different Versions of ChatGPT: A Transformative Breakthrough or the Next Black Box Challenge? | |
| dc.type | Article |










