-
摘要:
目的 探讨大语言模型(LLM)作为决策支持工具在急诊烧伤救治中的应用效果。 方法 该研究为横断面研究。2025年9月—2026年4月,浙江大学医学院附属第二医院急诊医学科收治100例符合入选标准的烧伤患者,其中男52例、女48例,年龄34.5(25.0,56.8)岁。以患者的出院小结为基础,将由该院2名具有副高级职称的急诊医学科医师与1名具有正高级职称的烧伤与创面修复科医师组成的专家组审核确定后的诊断和治疗方案作为“金标准”。由2名工作年限>5年的烧伤与创面修复科医师根据“金标准”,采用利克特量表对4种LLM即ChatGPT-5.5、豆包-2.0-Pro、千问-3.6-Max-Preview和Gemini 3.5 Flash和10名急诊医学科住院医师(EDJP)于2026年5月25—28日针对前述100例患者生成的诊断和治疗方案,通过盲法进行诊断准确性和治疗方案合理性评价。比较EDJP和4种LLM对100例患者的诊断准确性评分百分比与治疗方案合理性评分;在不同烧伤严重程度患者中,比较EDJP和4种LLM对患者的诊断准确性评分百分比;比较2名评价者诊断准确性评分、治疗方案合理性评分的一致性。 结果 豆包-2.0-Pro、Gemini 3.5 Flash、ChatGPT-5.5和千问-3.6-Max-Preview对100例患者的诊断准确性评分百分比分别为100%(100%,100%)、100%(100%,100%)、100%(50.0%,100%)和100%(50.0%,100%),均明显高于EDJP的100%(50.0%,100%),P<0.05。豆包-2.0-Pro、ChatGPT-5.5和Gemini 3.5 Flash对100例患者的治疗方案合理性评分分别为4.0(4.0,4.0)、4.0(4.0,4.0)、4.0(4.0,4.0)分,均明显高于EDJP的4.0(3.0,4.0)分(P<0.05);豆包-2.0-Pro和ChatGPT-5.5对100例患者的治疗方案合理性评分均明显高于千问-3.6-Max-Preview的4.0(3.0,4.0)分(P<0.05)。豆包-2.0-Pro、Gemini 3.5 Flash、ChatGPT-5.5对轻度烧伤患者的诊断准确性评分百分比均明显高于EDJP(P<0.05);豆包-2.0-Pro、千问-3.6-Max-Preview对中度烧伤患者的诊断准确性评分百分比均明显高于EDJP(P<0.05);EDJP和4种LLM在对重度和特重度烧伤患者的诊断准确性评分百分比方面比较,差异均无统计学意义(P>0.05)。2名评价者诊断准确性评分呈较好的一致性(组内相关系数为0.821,95%CI为0.746~0.876),治疗方案合理性评分呈中等一致性(组内相关系数为0.650,95%CI为0.522~0.749)。 结论 豆包-2.0-Pro、ChatGPT-5.5和Gemini 3.5 Flash在对急诊烧伤患者的诊断准确性和治疗方案合理性方面达到甚至超越了EDJP,初步具备作为烧伤临床第二意见工具的潜力。 Abstract:Objective To explore the application effect of large language models (LLM) as decision-support tools in emergency burn treatment. Methods This study was a cross-sectional study. A total of 100 burn patients meeting the inclusion criteria were admitted to the Department of Emergency Medicine of the Second Affiliated Hospital of Zhejiang University School of Medicine between September 2025 and April 2026, including 52 males and 48 females, aged 34.5 (25.0, 56.8) years. Based on patients' discharge summaries, the diagnoses and treatment plans reviewed and confirmed by an expert panel consisting of two associate senior emergency medicine department physicians and one senior burn and wound repair department physician from the hospital were defined as the gold standard. Two burn and wound repair department physicians with more than 5 years of working experience conducted blind evaluations on the diagnostic accuracy and rationality of treatment plans for all diagnoses and treatment plans generated by 4 LLMs (ChatGPT-5.5, Doubao-2.0-Pro, Qwen-3.6-Max-Preview, and Gemini 3.5 Flash) as well as 10 emergency medicine department junior physicians (EDJPs) for the aforementioned 100 patients form May 25-28,2026 using the Likert scale in accordance with the gold standard. The percentage of diagnostic accuracy scores and treatment plan rationality scores for the 100 patients were compared among EDJPs and the four LLMs; the percentage of diagnostic accuracy scores across patients with varying burn severities were also compared among EDJPs and the four LLMs; the inter-rater consistency of diagnostic accuracy scores and treatment plan rationality scores between the two evaluators was assessed. Results The percentages of diagnostic accuracy scores for the 100 patients by Doubao-2.0-Pro, Gemini 3.5 Flash, ChatGPT-5.5, and Qwen-3.6-Max-Preview were 100% (100%, 100%), 100% (100%, 100%), 100% (50.0%, 100%), and 100% (50.0%, 100%), respectively, which were all significantly higher than the 100% (50.0%, 100%) by EDJPs (P<0.05). The treatment plan rationality scores for the 100 patients by Doubao-2.0-Pro, ChatGPT-5.5 and Gemini 3.5 Flash were 4.0 (4.0, 4.0), 4.0 (4.0, 4.0), and 4.0 (4.0, 4.0), respectively, which were all significantly higher than the 4.0 (3.0, 4.0) by EDJPs (P<0.05). The treatment plan rationality scores for the 100 patients by Doubao-2.0-Pro and ChatGPT-5.5 were significantly higher than the 4.0 (3.0, 4.0) by Qwen-3.6-Max-Preview (P<0.05). The percentages of diagnostic accuracy scores for patients with mild burns by Doubao-2.0-Pro, Gemini 3.5 Flash, and ChatGPT-5.5 were significantly higher than that by EDJPs (P<0.05); the percentages of diagnostic accuracy scores for patients with moderate burns by Doubao-2.0-Pro and Qwen-3.6-Max-Preview were significantly higher than that by EDJPs (P<0.05). No statistically significant differences in percentages of diagnostic accuracy scores were found among EDJPs and the four LLMs for patients with severe and extremely severe burns (P>0.05). Good inter-rater consistency was observed for diagnostic accuracy scores between the two evaluators (with a interclass correlation coefficient of 0.821, with a 95% CI of 0.746-0.876), and moderate consistency was noted for treatment plan rationality scores (with a interclass correlation coefficient of 0.650, with a 95% CI of 0.522-0.749). Conclusions Doubao-2.0-Pro, ChatGPT-5.5, and Gemini 3.5 Flash demonstrate diagnostic accuracy and treatment plan rationality equivalent to or even superior to those of EDJPs for patients with emergency burns, showing preliminary potential to serve as burn clinical second-opinion tools. -
参考文献
(39) [1] ShaoJL,ZhuYQ,XieMS,et al.Assessing the global burns burden in the context of the WHO prevention and care strategy: insights from the global burden of disease study 2021[J].Nurs Health Sci,2026,28(2):e70329.DOI: 10.1111/nhs.70329. [2] LiuT,QuY,ChaiJ,et al.Epidemiology and first aid measures in pediatric burn patients in northern China during 2016-2020: a single-center retrospective study[J].Health Sci Rep,2024,7(7):e2218.DOI: 10.1002/hsr2.2218. [3] 房嫣,丁培杰,方洁,等.利奈唑胺在烧伤患者中的药代动力学研究[J].中华烧伤与创面修复杂志,2026,42(5):459-466.DOI: 10.3760/cma.j.cn501225-20251028-00447. [4] OzelM,YilmazS,TatliparmakAC,et al.Comparative analysis of burn injuries in toddler and preschool children: implications for triage and outcome assessment[J].Disaster Med Public Health Prep,2024,18:e143.DOI: 10.1017/dmp.2024.146. [5] BrodeurPG,BuckleyTA,KanjeeZ,et al.Performance of a large language model on the reasoning tasks of a physician[J].Science,2026,392(6797):524-527.DOI: 10.1126/science.adz4433. [6] 汪洋,吴嘉豪,张帆,等.大语言模型在医学检验技术教育中的能力评估[J].中华医学教育探索杂志,2025,24(11):1447-1453.DOI: 10.3760/cma.j.cn116021-20250516-02108. [7] OmarM,SofferS,AgbareiaR,et al.Sociodemographic biases in medical decision making by large language models[J].Nat Med,2025,31(6):1873-1881.DOI: 10.1038/s41591-025-03626-6. [8] WorkumJD,VolkersBWS,van de SandeD,et al.Comparative evaluation and performance of large language models on expert level critical care questions: a benchmark study[J].Crit Care,2025,29(1):72.DOI: 10.1186/s13054-025-05302-0. [9] ArtsiY,SorinV,GlicksbergBS,et al.Large language models in real-world clinical workflows: a systematic review of applications and implementation[J].Front Digit Health,2025,7:1659134.DOI: 10.3389/fdgth.2025.1659134. [10] KhasentinoJ,BelyaevaA,LiuX,et al.A personal health large language model for sleep and fitness coaching[J].Nat Med,2025,31(10):3394-3403.DOI: 10.1038/s41591-025-03888-0. [11] MasanneckL,SchmidtL,SeifertA,et al.Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: comparative study[J].J Med Internet Res,2024,26:e53297.DOI: 10.2196/53297. [12] MarcacciniG,SethI,LimB,et al.Management of burns: multi-center assessment comparing AI models and experienced plastic surgeons[J].J Clin Med,2025,14(9):3078.DOI: 10.3390/jcm14093078. [13] WeiJ,JiangS,YinT,et al.Performance evaluation of large language models in the diagnosis of emergency internal medicine diseases: a retrospective study[J].Front Public Health,2026,14:1780425.DOI: 10.3389/fpubh.2026.1780425. [14] AykutA,KarayılAR,YıldırımC,et al.Multimodal large language model versus emergency physicians for burn assessment: a prospective non-inferiority study[J].Scand J Trauma Resusc Emerg Med,2026,34(1):54.DOI: 10.1186/s13049-026-01577-6. [15] HagerP,JungmannF,HollandR,et al.Evaluation and mitigation of the limitations of large language models in clinical decision-making[J].Nat Med,2024,30(9):2613-2622.DOI: 10.1038/s41591-024-03097-1. [16] BediS,LiuY,Orr-EwingL,et al.Testing and evaluation of health care applications of large language models: a systematic review[J].JAMA,2025,333(4):319-328.DOI: 10.1001/jama.2024.21700. [17] KresoA,BobanZ,KabicS,et al.Using large language models as decision support tools in emergency ophthalmology[J].Int J Med Inform,2025,199:105886.DOI: 10.1016/j.ijmedinf.2025.105886. [18] TordjmanM,LiuZ,YuceM,et al.Comparative benchmarking of the DeepSeek large language model on medical tasks and clinical reasoning[J].Nat Med,2025,31(8):2550-2555.DOI: 10.1038/s41591-025-03726-3. [19] 于学忠,陆一鸣.急诊医学[M].2版.北京:人民卫生出版社,2021:458-462. [20] GongEJ,BangCS,LeeJJ,et al.Knowledge-practice performance gap in clinical large language models: systematic review of 39 benchmarks[J].J Med Internet Res,2025,27:e84120.DOI: 10.2196/84120. [21] 胡燕飞,王艾,张亚平,等.基于mT5大语言模型建立用于住院医师培训的胸部CT报告结论生成系统[J].中华医学教育探索杂志,2025,24(8):1016-1021.DOI: 10.3760/cma.j.cn116021-20240710-02089. [22] XuL,ZhaoW,HuangX.Diagnosis and triage performance of contemporary large language models on short clinical vignettes[J].J Med Syst,2025,49(1):141.DOI: 10.1007/s10916-025-02284-y. [23] PatelA,RuoffC,HelgesonSA,et al.Diagnostic performance of Large Language Models (LLMs) compared with physicians in sleep medicine[J].Sleep Med,2025,134:106677.DOI: 10.1016/j.sleep.2025.106677. [24] GuoY,MengX,YuE,et al.Development and prospective shadow evaluation of a domain-specific large language model for emergency neurological diagnosis[J].NPJ Digit Med,2026,9(1):470.DOI: 10.1038/s41746-026-02644-z. [25] RaoAS,EsmailKP,LeeRS,et al.Large language model performance and clinical reasoning tasks[J].JAMA Netw Open,2026,9(4):e264003.DOI: 10.1001/jamanetworkopen.2026.4003. [26] 张玉雯,白颖璐,张学敏,等.人工智能技术在脓毒症患者诊断与治疗中应用的研究进展[J].中华烧伤与创面修复杂志,2025,41(10):998-1003.DOI: 10.3760/cma.j.cn501225-20250708-00292. [27] ParkSH,SuhCH,LeeJH,et al.Minimum reporting items for clear evaluation of accuracy reports of large language models in healthcare (MI-CLEAR-LLM): 2025 Updates[J].Korean J Radiol,2025,26(12):1123-1132.DOI: 10.3348/kjr.2025.1522. [28] WilliamsG,RutundaS,NzabakiraF,et al.Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of GenAI in global health[J/OL].NPJ Digit Med,2026(2026-07-20)[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/42477479/.DOI:10.1038/s41746-026-02992-w.[published online ahead of print]. [29] WuX,HuangY,HeQ.A large language model improves clinicians' diagnostic performance in complex critical illness cases[J].Crit Care,2025,29(1):230.DOI: 10.1186/s13054-025-05468-7. [30] RowlandR,PonticorvoA,BaldadoM,et al.Burn wound classification model using spatial frequency-domain imaging and machine learning[J].J Biomed Opt,2019,24(5):1-9.DOI: 10.1117/1.JBO.24.5.056007. [31] ChauhanJ,GoyalP.BPBSAM: body part-specific burn severity assessment model[J].Burns,2020,46(6):1407-1423.DOI: 10.1016/j.burns.2020.03.007. [32] WuD,LuJ,XuD.Performances of five large language models in clinical decision-making for internal medicine: a comparative study[J].Digit Health,2026,12:20552076261433073.DOI: 10.1177/20552076261433073. [33] BoztasAE,GenisolI,PayzaAD,et al.ChatGPT-4o in pediatric burn care: expert review of its role in initial clinical decision-making[J].J Burn Care Res,2026,47(2):620-628.DOI: 10.1093/jbcr/iraf211. [34] JiS,XiaoS,XiaZ,et al.Consensus on the treatment of second-degree burn wounds (2024 edition)[J/OL].Burns Trauma,2024,12:tkad061[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/38343901/.DOI: 10.1093/burnst/tkad061. [35] LevineDM,TuwaniR,KompaB,et al.The diagnostic and triage accuracy of the GPT-3 artificial intelligence model: an observational study[J].Lancet Digit Health,2024,6(8):e555-e561.DOI: 10.1016/S2589-7500(24)00097-9. [36] WangL,TangK,WangY,et al.Advancements in artificial intelligence-driven diagnostic models for traditional Chinese medicine[J].Am J Chin Med,2025,53(3):647-673.DOI: 10.1142/S0192415X25500259. [37] YangJ,YangS,ShaoZ,et al.Large language models in emergency and critical care medicine: a comprehensive review of applications, challenges, and future directions[J/OL].Burns Trauma,2026,14:tkag026[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/42261384/.DOI: 10.1093/burnst/tkag026. [38] OmarM,SorinV,CollinsJD,et al.Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support[J].Commun Med (Lond),2025,5(1):330.DOI: 10.1038/s43856-025-01021-3. [39] MudumbaiSC,ChungP,ChenJQ,et al.Evaluating large language model performance in generating clinically relevant intensive care unit discharge summaries[J].A A Pract,2025,19(9):e02057.DOI: 10.1213/XAA.0000000000002057. -
Table 1. EDJP和4种大语言模型对不同严重程度烧伤患者的诊断准确性评分百分比比较[%,M(Q1,Q3)]
严重程度 例数 EDJP 豆包-2.0-Pro 千问-3.6-Max-Preview Gemini 3.5 Flash ChatGPT-5.5 F值 P值 轻度 49 100(50.0,100) 100(100,100)a 100(50.0,100) 100(100,100)ab 100(100,100)a 20.720 0.001 中度 11 50.0(50.0,50.0) 100(100,100)a 100(50.0,100)a 100(50.0,100)a 100(50.0,100) 14.180 0.007 重度 35 100(50.0,100) 100(100,100) 100(100,100) 100(100,100) 100(100,100) 6.000 0.199 特重度 5 100(100,100) 100(50.0,100) 100(100,100) 100(100,100) 100(100,100) 6.435 0.169 注:EDJP为急诊医学科住院医师;F值、P值为每种严重程度中EDJP和4种大语言模型总体比较所得;与EDJP比较,aP<0.05;与千问-3.6-Max-Preview比较,bP<0.05 -



下载: