留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

LLM作为决策支持工具在急诊烧伤救治中的应用效果

颜梦婷 魏金涛 蒋守银 吴攀 王新刚

颜梦婷, 魏金涛, 蒋守银, 等. LLM作为决策支持工具在急诊烧伤救治中的应用效果[J]. 中华烧伤与创面修复杂志, 2026, 42(9): 1-9. DOI: 10.3760/cma.j.cn501225-20260601-00223.
引用本文: 颜梦婷, 魏金涛, 蒋守银, 等. LLM作为决策支持工具在急诊烧伤救治中的应用效果[J]. 中华烧伤与创面修复杂志, 2026, 42(9): 1-9. DOI: 10.3760/cma.j.cn501225-20260601-00223.
Yan Mengting,Wei Jintao,Jiang Shouyin,et al.Application effect of LLMs as decision-support tools in emergency burn treatment[J].Chin J Burns Wounds,2026,42(9):1-9.DOI: 10.3760/cma.j.cn501225-20260601-00223.
Citation: Yan Mengting,Wei Jintao,Jiang Shouyin,et al.Application effect of LLMs as decision-support tools in emergency burn treatment[J].Chin J Burns Wounds,2026,42(9):1-9.DOI: 10.3760/cma.j.cn501225-20260601-00223.

LLM作为决策支持工具在急诊烧伤救治中的应用效果

doi: 10.3760/cma.j.cn501225-20260601-00223
基金项目: 

浙江省全省重点实验室项目 2024E10076

详细信息
    通讯作者:

    王新刚,Email:wangxingang8157@zju.edu.cn

Application effect of LLMs as decision-support tools in emergency burn treatment

Funds: 

Zhejiang Provincial Key Laboratory Project 2024E10076

More Information
  • 摘要:   目的  探讨大语言模型(LLM)作为决策支持工具在急诊烧伤救治中的应用效果。  方法  该研究为横断面研究。2025年9月—2026年4月,浙江大学医学院附属第二医院急诊医学科收治100例符合入选标准的烧伤患者,其中男52例、女48例,年龄34.5(25.0,56.8)岁。以患者的出院小结为基础,将由该院2名具有副高级职称的急诊医学科医师与1名具有正高级职称的烧伤与创面修复科医师组成的专家组审核确定后的诊断和治疗方案作为“金标准”。由2名工作年限>5年的烧伤与创面修复科医师根据“金标准”,采用利克特量表对4种LLM即ChatGPT-5.5、豆包-2.0-Pro、千问-3.6-Max-Preview和Gemini 3.5 Flash和10名急诊医学科住院医师(EDJP)于2026年5月25—28日针对前述100例患者生成的诊断和治疗方案,通过盲法进行诊断准确性和治疗方案合理性评价。比较EDJP和4种LLM对100例患者的诊断准确性评分百分比与治疗方案合理性评分;在不同烧伤严重程度患者中,比较EDJP和4种LLM对患者的诊断准确性评分百分比;比较2名评价者诊断准确性评分、治疗方案合理性评分的一致性。  结果  豆包-2.0-Pro、Gemini 3.5 Flash、ChatGPT-5.5和千问-3.6-Max-Preview对100例患者的诊断准确性评分百分比分别为100%(100%,100%)、100%(100%,100%)、100%(50.0%,100%)和100%(50.0%,100%),均明显高于EDJP的100%(50.0%,100%),P<0.05。豆包-2.0-Pro、ChatGPT-5.5和Gemini 3.5 Flash对100例患者的治疗方案合理性评分分别为4.0(4.0,4.0)、4.0(4.0,4.0)、4.0(4.0,4.0)分,均明显高于EDJP的4.0(3.0,4.0)分(P<0.05);豆包-2.0-Pro和ChatGPT-5.5对100例患者的治疗方案合理性评分均明显高于千问-3.6-Max-Preview的4.0(3.0,4.0)分(P<0.05)。豆包-2.0-Pro、Gemini 3.5 Flash、ChatGPT-5.5对轻度烧伤患者的诊断准确性评分百分比均明显高于EDJP(P<0.05);豆包-2.0-Pro、千问-3.6-Max-Preview对中度烧伤患者的诊断准确性评分百分比均明显高于EDJP(P<0.05);EDJP和4种LLM在对重度和特重度烧伤患者的诊断准确性评分百分比方面比较,差异均无统计学意义(P>0.05)。2名评价者诊断准确性评分呈较好的一致性(组内相关系数为0.821,95%CI为0.746~0.876),治疗方案合理性评分呈中等一致性(组内相关系数为0.650,95%CI为0.522~0.749)。  结论  豆包-2.0-Pro、ChatGPT-5.5和Gemini 3.5 Flash在对急诊烧伤患者的诊断准确性和治疗方案合理性方面达到甚至超越了EDJP,初步具备作为烧伤临床第二意见工具的潜力。

     

  • 参考文献(39)

    [1] ShaoJL,ZhuYQ,XieMS,et al.Assessing the global burns burden in the context of the WHO prevention and care strategy: insights from the global burden of disease study 2021[J].Nurs Health Sci,2026,28(2):e70329.DOI: 10.1111/nhs.70329.
    [2] LiuT,QuY,ChaiJ,et al.Epidemiology and first aid measures in pediatric burn patients in northern China during 2016-2020: a single-center retrospective study[J].Health Sci Rep,2024,7(7):e2218.DOI: 10.1002/hsr2.2218.
    [3] 房嫣,丁培杰,方洁,等.利奈唑胺在烧伤患者中的药代动力学研究[J].中华烧伤与创面修复杂志,2026,42(5):459-466.DOI: 10.3760/cma.j.cn501225-20251028-00447.
    [4] OzelM,YilmazS,TatliparmakAC,et al.Comparative analysis of burn injuries in toddler and preschool children: implications for triage and outcome assessment[J].Disaster Med Public Health Prep,2024,18:e143.DOI: 10.1017/dmp.2024.146.
    [5] BrodeurPG,BuckleyTA,KanjeeZ,et al.Performance of a large language model on the reasoning tasks of a physician[J].Science,2026,392(6797):524-527.DOI: 10.1126/science.adz4433.
    [6] 汪洋,吴嘉豪,张帆,等.大语言模型在医学检验技术教育中的能力评估[J].中华医学教育探索杂志,2025,24(11):1447-1453.DOI: 10.3760/cma.j.cn116021-20250516-02108.
    [7] OmarM,SofferS,AgbareiaR,et al.Sociodemographic biases in medical decision making by large language models[J].Nat Med,2025,31(6):1873-1881.DOI: 10.1038/s41591-025-03626-6.
    [8] WorkumJD,VolkersBWS,van de SandeD,et al.Comparative evaluation and performance of large language models on expert level critical care questions: a benchmark study[J].Crit Care,2025,29(1):72.DOI: 10.1186/s13054-025-05302-0.
    [9] ArtsiY,SorinV,GlicksbergBS,et al.Large language models in real-world clinical workflows: a systematic review of applications and implementation[J].Front Digit Health,2025,7:1659134.DOI: 10.3389/fdgth.2025.1659134.
    [10] KhasentinoJ,BelyaevaA,LiuX,et al.A personal health large language model for sleep and fitness coaching[J].Nat Med,2025,31(10):3394-3403.DOI: 10.1038/s41591-025-03888-0.
    [11] MasanneckL,SchmidtL,SeifertA,et al.Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: comparative study[J].J Med Internet Res,2024,26:e53297.DOI: 10.2196/53297.
    [12] MarcacciniG,SethI,LimB,et al.Management of burns: multi-center assessment comparing AI models and experienced plastic surgeons[J].J Clin Med,2025,14(9):3078.DOI: 10.3390/jcm14093078.
    [13] WeiJ,JiangS,YinT,et al.Performance evaluation of large language models in the diagnosis of emergency internal medicine diseases: a retrospective study[J].Front Public Health,2026,14:1780425.DOI: 10.3389/fpubh.2026.1780425.
    [14] AykutA,KarayılAR,YıldırımC,et al.Multimodal large language model versus emergency physicians for burn assessment: a prospective non-inferiority study[J].Scand J Trauma Resusc Emerg Med,2026,34(1):54.DOI: 10.1186/s13049-026-01577-6.
    [15] HagerP,JungmannF,HollandR,et al.Evaluation and mitigation of the limitations of large language models in clinical decision-making[J].Nat Med,2024,30(9):2613-2622.DOI: 10.1038/s41591-024-03097-1.
    [16] BediS,LiuY,Orr-EwingL,et al.Testing and evaluation of health care applications of large language models: a systematic review[J].JAMA,2025,333(4):319-328.DOI: 10.1001/jama.2024.21700.
    [17] KresoA,BobanZ,KabicS,et al.Using large language models as decision support tools in emergency ophthalmology[J].Int J Med Inform,2025,199:105886.DOI: 10.1016/j.ijmedinf.2025.105886.
    [18] TordjmanM,LiuZ,YuceM,et al.Comparative benchmarking of the DeepSeek large language model on medical tasks and clinical reasoning[J].Nat Med,2025,31(8):2550-2555.DOI: 10.1038/s41591-025-03726-3.
    [19] 于学忠,陆一鸣.急诊医学[M].2版.北京:人民卫生出版社,2021:458-462.
    [20] GongEJ,BangCS,LeeJJ,et al.Knowledge-practice performance gap in clinical large language models: systematic review of 39 benchmarks[J].J Med Internet Res,2025,27:e84120.DOI: 10.2196/84120.
    [21] 胡燕飞,王艾,张亚平,等.基于mT5大语言模型建立用于住院医师培训的胸部CT报告结论生成系统[J].中华医学教育探索杂志,2025,24(8):1016-1021.DOI: 10.3760/cma.j.cn116021-20240710-02089.
    [22] XuL,ZhaoW,HuangX.Diagnosis and triage performance of contemporary large language models on short clinical vignettes[J].J Med Syst,2025,49(1):141.DOI: 10.1007/s10916-025-02284-y.
    [23] PatelA,RuoffC,HelgesonSA,et al.Diagnostic performance of Large Language Models (LLMs) compared with physicians in sleep medicine[J].Sleep Med,2025,134:106677.DOI: 10.1016/j.sleep.2025.106677.
    [24] GuoY,MengX,YuE,et al.Development and prospective shadow evaluation of a domain-specific large language model for emergency neurological diagnosis[J].NPJ Digit Med,2026,9(1):470.DOI: 10.1038/s41746-026-02644-z.
    [25] RaoAS,EsmailKP,LeeRS,et al.Large language model performance and clinical reasoning tasks[J].JAMA Netw Open,2026,9(4):e264003.DOI: 10.1001/jamanetworkopen.2026.4003.
    [26] 张玉雯,白颖璐,张学敏,等.人工智能技术在脓毒症患者诊断与治疗中应用的研究进展[J].中华烧伤与创面修复杂志,2025,41(10):998-1003.DOI: 10.3760/cma.j.cn501225-20250708-00292.
    [27] ParkSH,SuhCH,LeeJH,et al.Minimum reporting items for clear evaluation of accuracy reports of large language models in healthcare (MI-CLEAR-LLM): 2025 Updates[J].Korean J Radiol,2025,26(12):1123-1132.DOI: 10.3348/kjr.2025.1522.
    [28] WilliamsG,RutundaS,NzabakiraF,et al.Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of GenAI in global health[J/OL].NPJ Digit Med,2026(2026-07-20)[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/42477479/.DOI:10.1038/s41746-026-02992-w.[published online ahead of print].
    [29] WuX,HuangY,HeQ.A large language model improves clinicians' diagnostic performance in complex critical illness cases[J].Crit Care,2025,29(1):230.DOI: 10.1186/s13054-025-05468-7.
    [30] RowlandR,PonticorvoA,BaldadoM,et al.Burn wound classification model using spatial frequency-domain imaging and machine learning[J].J Biomed Opt,2019,24(5):1-9.DOI: 10.1117/1.JBO.24.5.056007.
    [31] ChauhanJ,GoyalP.BPBSAM: body part-specific burn severity assessment model[J].Burns,2020,46(6):1407-1423.DOI: 10.1016/j.burns.2020.03.007.
    [32] WuD,LuJ,XuD.Performances of five large language models in clinical decision-making for internal medicine: a comparative study[J].Digit Health,2026,12:20552076261433073.DOI: 10.1177/20552076261433073.
    [33] BoztasAE,GenisolI,PayzaAD,et al.ChatGPT-4o in pediatric burn care: expert review of its role in initial clinical decision-making[J].J Burn Care Res,2026,47(2):620-628.DOI: 10.1093/jbcr/iraf211.
    [34] JiS,XiaoS,XiaZ,et al.Consensus on the treatment of second-degree burn wounds (2024 edition)[J/OL].Burns Trauma,2024,12:tkad061[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/38343901/.DOI: 10.1093/burnst/tkad061.
    [35] LevineDM,TuwaniR,KompaB,et al.The diagnostic and triage accuracy of the GPT-3 artificial intelligence model: an observational study[J].Lancet Digit Health,2024,6(8):e555-e561.DOI: 10.1016/S2589-7500(24)00097-9.
    [36] WangL,TangK,WangY,et al.Advancements in artificial intelligence-driven diagnostic models for traditional Chinese medicine[J].Am J Chin Med,2025,53(3):647-673.DOI: 10.1142/S0192415X25500259.
    [37] YangJ,YangS,ShaoZ,et al.Large language models in emergency and critical care medicine: a comprehensive review of applications, challenges, and future directions[J/OL].Burns Trauma,2026,14:tkag026[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/42261384/.DOI: 10.1093/burnst/tkag026.
    [38] OmarM,SorinV,CollinsJD,et al.Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support[J].Commun Med (Lond),2025,5(1):330.DOI: 10.1038/s43856-025-01021-3.
    [39] MudumbaiSC,ChungP,ChenJQ,et al.Evaluating large language model performance in generating clinically relevant intensive care unit discharge summaries[J].A A Pract,2025,19(9):e02057.DOI: 10.1213/XAA.0000000000002057.
  • 图  1  EDJP和4种大语言模型对100例烧伤患者的诊断准确性评分分布

    注:EDJP为急诊医学科住院医师;绿色为0分,浅黄色为1分,蓝色为2分

    图  2  EDJP和4种大语言模型对100例烧伤患者的治疗方案合理性评分分布

    注:EDJP为急诊医学科住院医师;淡紫色为1分,黄褐色为2分,浅米色为3分,湖蓝色为4分

    Table  1.   EDJP和4种大语言模型对不同严重程度烧伤患者的诊断准确性评分百分比比较[%,MQ1,Q3)]

    严重程度例数EDJP豆包-2.0-Pro千问-3.6-Max-PreviewGemini 3.5 FlashChatGPT-5.5FP
    轻度49100(50.0,100)100(100,100)a100(50.0,100)100(100,100)ab100(100,100)a20.7200.001
    中度1150.0(50.0,50.0)100(100,100)a100(50.0,100)a100(50.0,100)a100(50.0,100)14.1800.007
    重度35100(50.0,100)100(100,100)100(100,100)100(100,100)100(100,100)6.0000.199
    特重度5100(100,100)100(50.0,100)100(100,100)100(100,100)100(100,100)6.4350.169
    注:EDJP为急诊医学科住院医师;F值、P值为每种严重程度中EDJP和4种大语言模型总体比较所得;与EDJP比较,aP<0.05;与千问-3.6-Max-Preview比较,bP<0.05
    下载: 导出CSV
  • 加载中
图(3) / 表(1)
计量
  • 文章访问数:  1
  • HTML全文浏览量:  0
  • PDF下载量:  0
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-06-01
  • 网络出版日期:  2026-08-27

目录

    /

    返回文章
    返回