Laboratory Medicine Decision Support-Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek-LLM Decision Support in Laboratory Medicine. [PDF]
Ulutaş KT, Pekmezci A.
europepmc +1 more source
Development of a Framework for Evaluating Large Language Model Safety and Reliability: a Proof-of-Concept Evaluation. [PDF]
Liu F +6 more
europepmc +1 more source
Temporal Analysis of Patient-Centered Sentiment in Clinical Notes for Patients With Mental Health Conditions: Retrospective Cohort Study. [PDF]
Morsy A, Peters CJ, Miller L, Zirikly A.
europepmc +1 more source
Evaluating AI-generated patient education materials for endometrial cancer surgery: a comparative analysis of response quality, reliability, and readability between ChatGPT and DeepSeek models. [PDF]
Tian L +9 more
europepmc +1 more source
Large Language Model Performance on Multistep Clinical Cases: Comparative Study Across Question and Case Levels. [PDF]
Cha J, Zhao Y, Zong H.
europepmc +1 more source
Agreement in Staging and Treatment Recommendations Among Clinicians, Text-Only Clinicians, DeepSeek-V3, and ChatGPT-4o for Nasopharyngeal Carcinoma Patients. [PDF]
Lin S +5 more
europepmc +1 more source
Comprehensive Evaluation of Large Language Models on Four Core Medical School Courses: A Cross-Sectional Comparative Study. [PDF]
Zhang K, Yang W, Zheng W.
europepmc +1 more source
Assessing Guideline Knowledge Alignment of Large Language Models in Ophthalmology: A Preclinical Benchmarking Study on KLEx Evidence-Based Guidelines. [PDF]
Lu H +7 more
europepmc +1 more source
Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study. [PDF]
Zhang J +7 more
europepmc +1 more source
How AI responds to common HIV/AIDS questions: ChatGPT versus DeepSeek. [PDF]
Huang H +10 more
europepmc +1 more source

