Turn off MathJax
Article Contents
Yan Mengting,Wei Jintao,Jiang Shouyin,et al.Application effect of LLMs as decision-support tools in emergency burn treatment[J].Chin J Burns Wounds,2026,42(9):1-9.DOI: 10.3760/cma.j.cn501225-20260601-00223.
Citation: Yan Mengting,Wei Jintao,Jiang Shouyin,et al.Application effect of LLMs as decision-support tools in emergency burn treatment[J].Chin J Burns Wounds,2026,42(9):1-9.DOI: 10.3760/cma.j.cn501225-20260601-00223.

Application effect of LLMs as decision-support tools in emergency burn treatment

doi: 10.3760/cma.j.cn501225-20260601-00223
Funds:

Zhejiang Provincial Key Laboratory Project 2024E10076

More Information
  •   Objective  To explore the application effect of large language models (LLM) as decision-support tools in emergency burn treatment.  Methods  This study was a cross-sectional study. A total of 100 burn patients meeting the inclusion criteria were admitted to the Department of Emergency Medicine of the Second Affiliated Hospital of Zhejiang University School of Medicine between September 2025 and April 2026, including 52 males and 48 females, aged 34.5 (25.0, 56.8) years. Based on patients' discharge summaries, the diagnoses and treatment plans reviewed and confirmed by an expert panel consisting of two associate senior emergency medicine department physicians and one senior burn and wound repair department physician from the hospital were defined as the gold standard. Two burn and wound repair department physicians with more than 5 years of working experience conducted blind evaluations on the diagnostic accuracy and rationality of treatment plans for all diagnoses and treatment plans generated by 4 LLMs (ChatGPT-5.5, Doubao-2.0-Pro, Qwen-3.6-Max-Preview, and Gemini 3.5 Flash) as well as 10 emergency medicine department junior physicians (EDJPs) for the aforementioned 100 patients form May 25-28,2026 using the Likert scale in accordance with the gold standard. The percentage of diagnostic accuracy scores and treatment plan rationality scores for the 100 patients were compared among EDJPs and the four LLMs; the percentage of diagnostic accuracy scores across patients with varying burn severities were also compared among EDJPs and the four LLMs; the inter-rater consistency of diagnostic accuracy scores and treatment plan rationality scores between the two evaluators was assessed.  Results  The percentages of diagnostic accuracy scores for the 100 patients by Doubao-2.0-Pro, Gemini 3.5 Flash, ChatGPT-5.5, and Qwen-3.6-Max-Preview were 100% (100%, 100%), 100% (100%, 100%), 100% (50.0%, 100%), and 100% (50.0%, 100%), respectively, which were all significantly higher than the 100% (50.0%, 100%) by EDJPs (P<0.05). The treatment plan rationality scores for the 100 patients by Doubao-2.0-Pro, ChatGPT-5.5 and Gemini 3.5 Flash were 4.0 (4.0, 4.0), 4.0 (4.0, 4.0), and 4.0 (4.0, 4.0), respectively, which were all significantly higher than the 4.0 (3.0, 4.0) by EDJPs (P<0.05). The treatment plan rationality scores for the 100 patients by Doubao-2.0-Pro and ChatGPT-5.5 were significantly higher than the 4.0 (3.0, 4.0) by Qwen-3.6-Max-Preview (P<0.05). The percentages of diagnostic accuracy scores for patients with mild burns by Doubao-2.0-Pro, Gemini 3.5 Flash, and ChatGPT-5.5 were significantly higher than that by EDJPs (P<0.05); the percentages of diagnostic accuracy scores for patients with moderate burns by Doubao-2.0-Pro and Qwen-3.6-Max-Preview were significantly higher than that by EDJPs (P<0.05). No statistically significant differences in percentages of diagnostic accuracy scores were found among EDJPs and the four LLMs for patients with severe and extremely severe burns (P>0.05). Good inter-rater consistency was observed for diagnostic accuracy scores between the two evaluators (with a interclass correlation coefficient of 0.821, with a 95% CI of 0.746-0.876), and moderate consistency was noted for treatment plan rationality scores (with a interclass correlation coefficient of 0.650, with a 95% CI of 0.522-0.749).  Conclusions  Doubao-2.0-Pro, ChatGPT-5.5, and Gemini 3.5 Flash demonstrate diagnostic accuracy and treatment plan rationality equivalent to or even superior to those of EDJPs for patients with emergency burns, showing preliminary potential to serve as burn clinical second-opinion tools.

     

  • loading
  • [1]
    ShaoJL,ZhuYQ,XieMS,et al.Assessing the global burns burden in the context of the WHO prevention and care strategy: insights from the global burden of disease study 2021[J].Nurs Health Sci,2026,28(2):e70329.DOI: 10.1111/nhs.70329.
    [2]
    LiuT,QuY,ChaiJ,et al.Epidemiology and first aid measures in pediatric burn patients in northern China during 2016-2020: a single-center retrospective study[J].Health Sci Rep,2024,7(7):e2218.DOI: 10.1002/hsr2.2218.
    [3]
    房嫣,丁培杰,方洁,等.利奈唑胺在烧伤患者中的药代动力学研究[J].中华烧伤与创面修复杂志,2026,42(5):459-466.DOI: 10.3760/cma.j.cn501225-20251028-00447.
    [4]
    OzelM,YilmazS,TatliparmakAC,et al.Comparative analysis of burn injuries in toddler and preschool children: implications for triage and outcome assessment[J].Disaster Med Public Health Prep,2024,18:e143.DOI: 10.1017/dmp.2024.146.
    [5]
    BrodeurPG,BuckleyTA,KanjeeZ,et al.Performance of a large language model on the reasoning tasks of a physician[J].Science,2026,392(6797):524-527.DOI: 10.1126/science.adz4433.
    [6]
    汪洋,吴嘉豪,张帆,等.大语言模型在医学检验技术教育中的能力评估[J].中华医学教育探索杂志,2025,24(11):1447-1453.DOI: 10.3760/cma.j.cn116021-20250516-02108.
    [7]
    OmarM,SofferS,AgbareiaR,et al.Sociodemographic biases in medical decision making by large language models[J].Nat Med,2025,31(6):1873-1881.DOI: 10.1038/s41591-025-03626-6.
    [8]
    WorkumJD,VolkersBWS,van de SandeD,et al.Comparative evaluation and performance of large language models on expert level critical care questions: a benchmark study[J].Crit Care,2025,29(1):72.DOI: 10.1186/s13054-025-05302-0.
    [9]
    ArtsiY,SorinV,GlicksbergBS,et al.Large language models in real-world clinical workflows: a systematic review of applications and implementation[J].Front Digit Health,2025,7:1659134.DOI: 10.3389/fdgth.2025.1659134.
    [10]
    KhasentinoJ,BelyaevaA,LiuX,et al.A personal health large language model for sleep and fitness coaching[J].Nat Med,2025,31(10):3394-3403.DOI: 10.1038/s41591-025-03888-0.
    [11]
    MasanneckL,SchmidtL,SeifertA,et al.Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: comparative study[J].J Med Internet Res,2024,26:e53297.DOI: 10.2196/53297.
    [12]
    MarcacciniG,SethI,LimB,et al.Management of burns: multi-center assessment comparing AI models and experienced plastic surgeons[J].J Clin Med,2025,14(9):3078.DOI: 10.3390/jcm14093078.
    [13]
    WeiJ,JiangS,YinT,et al.Performance evaluation of large language models in the diagnosis of emergency internal medicine diseases: a retrospective study[J].Front Public Health,2026,14:1780425.DOI: 10.3389/fpubh.2026.1780425.
    [14]
    AykutA,KarayılAR,YıldırımC,et al.Multimodal large language model versus emergency physicians for burn assessment: a prospective non-inferiority study[J].Scand J Trauma Resusc Emerg Med,2026,34(1):54.DOI: 10.1186/s13049-026-01577-6.
    [15]
    HagerP,JungmannF,HollandR,et al.Evaluation and mitigation of the limitations of large language models in clinical decision-making[J].Nat Med,2024,30(9):2613-2622.DOI: 10.1038/s41591-024-03097-1.
    [16]
    BediS,LiuY,Orr-EwingL,et al.Testing and evaluation of health care applications of large language models: a systematic review[J].JAMA,2025,333(4):319-328.DOI: 10.1001/jama.2024.21700.
    [17]
    KresoA,BobanZ,KabicS,et al.Using large language models as decision support tools in emergency ophthalmology[J].Int J Med Inform,2025,199:105886.DOI: 10.1016/j.ijmedinf.2025.105886.
    [18]
    TordjmanM,LiuZ,YuceM,et al.Comparative benchmarking of the DeepSeek large language model on medical tasks and clinical reasoning[J].Nat Med,2025,31(8):2550-2555.DOI: 10.1038/s41591-025-03726-3.
    [19]
    于学忠,陆一鸣.急诊医学[M].2版.北京:人民卫生出版社,2021:458-462.
    [20]
    GongEJ,BangCS,LeeJJ,et al.Knowledge-practice performance gap in clinical large language models: systematic review of 39 benchmarks[J].J Med Internet Res,2025,27:e84120.DOI: 10.2196/84120.
    [21]
    胡燕飞,王艾,张亚平,等.基于mT5大语言模型建立用于住院医师培训的胸部CT报告结论生成系统[J].中华医学教育探索杂志,2025,24(8):1016-1021.DOI: 10.3760/cma.j.cn116021-20240710-02089.
    [22]
    XuL,ZhaoW,HuangX.Diagnosis and triage performance of contemporary large language models on short clinical vignettes[J].J Med Syst,2025,49(1):141.DOI: 10.1007/s10916-025-02284-y.
    [23]
    PatelA,RuoffC,HelgesonSA,et al.Diagnostic performance of Large Language Models (LLMs) compared with physicians in sleep medicine[J].Sleep Med,2025,134:106677.DOI: 10.1016/j.sleep.2025.106677.
    [24]
    GuoY,MengX,YuE,et al.Development and prospective shadow evaluation of a domain-specific large language model for emergency neurological diagnosis[J].NPJ Digit Med,2026,9(1):470.DOI: 10.1038/s41746-026-02644-z.
    [25]
    RaoAS,EsmailKP,LeeRS,et al.Large language model performance and clinical reasoning tasks[J].JAMA Netw Open,2026,9(4):e264003.DOI: 10.1001/jamanetworkopen.2026.4003.
    [26]
    张玉雯,白颖璐,张学敏,等.人工智能技术在脓毒症患者诊断与治疗中应用的研究进展[J].中华烧伤与创面修复杂志,2025,41(10):998-1003.DOI: 10.3760/cma.j.cn501225-20250708-00292.
    [27]
    ParkSH,SuhCH,LeeJH,et al.Minimum reporting items for clear evaluation of accuracy reports of large language models in healthcare (MI-CLEAR-LLM): 2025 Updates[J].Korean J Radiol,2025,26(12):1123-1132.DOI: 10.3348/kjr.2025.1522.
    [28]
    WilliamsG,RutundaS,NzabakiraF,et al.Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of GenAI in global health[J/OL].NPJ Digit Med,2026(2026-07-20)[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/42477479/.DOI:10.1038/s41746-026-02992-w.[published online ahead of print].
    [29]
    WuX,HuangY,HeQ.A large language model improves clinicians' diagnostic performance in complex critical illness cases[J].Crit Care,2025,29(1):230.DOI: 10.1186/s13054-025-05468-7.
    [30]
    RowlandR,PonticorvoA,BaldadoM,et al.Burn wound classification model using spatial frequency-domain imaging and machine learning[J].J Biomed Opt,2019,24(5):1-9.DOI: 10.1117/1.JBO.24.5.056007.
    [31]
    ChauhanJ,GoyalP.BPBSAM: body part-specific burn severity assessment model[J].Burns,2020,46(6):1407-1423.DOI: 10.1016/j.burns.2020.03.007.
    [32]
    WuD,LuJ,XuD.Performances of five large language models in clinical decision-making for internal medicine: a comparative study[J].Digit Health,2026,12:20552076261433073.DOI: 10.1177/20552076261433073.
    [33]
    BoztasAE,GenisolI,PayzaAD,et al.ChatGPT-4o in pediatric burn care: expert review of its role in initial clinical decision-making[J].J Burn Care Res,2026,47(2):620-628.DOI: 10.1093/jbcr/iraf211.
    [34]
    JiS,XiaoS,XiaZ,et al.Consensus on the treatment of second-degree burn wounds (2024 edition)[J/OL].Burns Trauma,2024,12:tkad061[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/38343901/.DOI: 10.1093/burnst/tkad061.
    [35]
    LevineDM,TuwaniR,KompaB,et al.The diagnostic and triage accuracy of the GPT-3 artificial intelligence model: an observational study[J].Lancet Digit Health,2024,6(8):e555-e561.DOI: 10.1016/S2589-7500(24)00097-9.
    [36]
    WangL,TangK,WangY,et al.Advancements in artificial intelligence-driven diagnostic models for traditional Chinese medicine[J].Am J Chin Med,2025,53(3):647-673.DOI: 10.1142/S0192415X25500259.
    [37]
    YangJ,YangS,ShaoZ,et al.Large language models in emergency and critical care medicine: a comprehensive review of applications, challenges, and future directions[J/OL].Burns Trauma,2026,14:tkag026[2026-08-10].https://pubmed.ncbi.nlm.nih.gov/42261384/.DOI: 10.1093/burnst/tkag026.
    [38]
    OmarM,SorinV,CollinsJD,et al.Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support[J].Commun Med (Lond),2025,5(1):330.DOI: 10.1038/s43856-025-01021-3.
    [39]
    MudumbaiSC,ChungP,ChenJQ,et al.Evaluating large language model performance in generating clinically relevant intensive care unit discharge summaries[J].A A Pract,2025,19(9):e02057.DOI: 10.1213/XAA.0000000000002057.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(3)  / Tables(1)

    Article Metrics

    Article views (34) PDF downloads(3) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return