Other Titles

Development and Evaluation of an RAG-Based Nursing Documentation Chatbot [Poster Title]

Abstract

Background: Accurate nursing documentation is essential for communicating patient status and ensuring patient safety. However, heavy workloads and non-standard terminology often lead to omissions and errors. Large language models (LLMs) show promise in supporting nursing documentation, but hallucinations raise safety concerns. To mitigate these risks, we adopted a retrieval-augmented generation (RAG) approach that grounds responses in standard nursing terminologies [1]. This study aimed to develop and evaluate an RAG-based chatbot that supports nursing documentation using standard terminology and procedures.

Methods: From January 2022 to May 2024, we analyzed 70,926 unstructured nursing narratives from 76 patients with myocardial infarction who underwent percutaneous coronary intervention (PCI) at a tertiary hospital. Based on prior literature, only records from two days before to three days after PCI were included in the analysis[2,3]. The nursing records were segmented into meaningful units and used to construct a PCI nursing knowledge base grounded in NANDA–NIC–NOC (NNN), with each nursing statement linked to a corresponding SNOMED CT concept. This knowledge base served as the retrieval source for an RAG-assisted chatbot. Seventeen PCI clinical scenarios were submitted with identical prompts to three LLMs (Meta-Llama-3.1-8B, Mixtral-8x7B, and Qwen2.5-14B) to generate nursing documentation suggestions. Four nurses with at least two years of clinical experience rated each output on a 5-point Likert scale for accuracy, completeness, readability, safety, and human-centeredness[4]. Scores from the first run (run 1) were used to compare models, and the mean score across three runs was used to assess consistency.

Results: Nursing documentation most frequently focused on patient education, notification of physicians, vital sign monitoring, and confirmation of chest pain status. Across all scenarios, Qwen2.5-14B achieved the highest mean scores (accuracy: 4.69; completeness: 4.46; readability: 4.51; safety: 4.76; human-centeredness: 4.76/5), while Meta-Llama-3.1-8B and Mixtral-8x7B also scored above 4.0 in most domains. Completeness was relatively lower than the other domains for all models.

Conclusion: The RAG-based LLM chatbot using a standardized PCI nursing knowledge base generated nursing record suggestions with high accuracy and safety, primarily emphasizing patient monitoring and related interventions, which are central priorities in PCI nursing care.

Notes

This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Republic of Korea government (Ministry of Science and ICT) (RS-2024-00354718).

References:

[1] Silva JF, Almeida JR, Matos S. Extraction of family history information from clinical notes: deep learning and heuristics approach. JMIR Med Inform. 2020;8(12):e22898. doi:10.2196/22898. PMID: 33372893; PMCID: PMC7803476.

[2] Tamis-Holland JE, Abbott JD, Al-Azizi K, Barman N, Bortnick AE, Cohen MG, et al. SCAI expert consensus statement on the management of patients with STEMI referred for primary PCI. J Soc Cardiovasc Angiogr Interv. 2024;3(11):102294.

[3] Canadian Council of Cardiovascular Nurses. Cardiovascular nursing practice standards. Ottawa, ON: Canadian Council of Cardiovascular Nurses; 2023. Available from: https://www.cccn.ca/.

[4] Qiang S, Zhang H, Liao Y, Zhang Y, Gu Y, Wang Y, et al. Application of large language models in stroke rehabilitation health education: 2-phase study. J Med Internet Res. 2025;27:e73226. doi:10.2196/73226. PMID: 40694436; PMCID: PMC12306586.

Description

This study developed and evaluated a retrieval-augmented generation (RAG) chatbot to support PCI nursing documentation using standard terminology. A NANDA–NIC–NOC–SNOMED CT knowledge base was built from 70,926 nursing narratives of 76 PCI patients at a tertiary hospital. Three large language models generated documentation suggestions, which four nurses rated on a 5-point scale. All models scored above 4.0 in most domains, with Qwen2.5-14B showing the highest overall perf

Author Details

Sumi Sung, PhD; Hyeyoung Lee, Doctoral Student; Wooje Sung, PhD; Do Yeon Kim, BSN; Jung Eun Hong, PhD; Yewon Lee, PhD; Seunghee Lee, PhD 

Additional author details in poster.

Sigma Membership

Non-member

Type

Poster

Format Type

Text-based Document

Study Design/Type

Other

Research Approach

Other

Keywords:

Acute Care, Implementation Science, Academic-Clinical Partnership, Emerging Technologies, Language Models, Chatbots, Natural Language Processing, Documentation, Nursing Records

Conference Name

37th International Nursing Research Congress

Conference Host

Sigma Theta Tau International

Conference Location

Toronto, Ontario, Canada

Conference Year

2026

Rights Holder

All rights reserved by the author(s) and/or publisher(s) listed in this item record unless relinquished in whole or part by a rights notation or a Creative Commons License present in this item record. All permission requests should be directed accordingly and not to the Sigma Repository. All submitting authors or publishers have affirmed that when using material in their work where they do not own copyright, they have obtained permission of the copyright holder prior to submission and the rights holder has been acknowledged as necessary.

Review Type

Abstract Review Only: Reviewed by Event Host

Acquisition

Proxy-submission

Date of Issue

2026-08-27

Click on the above link to access the poster.

Share

COinS
 

Development and Evaluation of an LLM Based Nursing Documentation

Toronto, Ontario, Canada

Background: Accurate nursing documentation is essential for communicating patient status and ensuring patient safety. However, heavy workloads and non-standard terminology often lead to omissions and errors. Large language models (LLMs) show promise in supporting nursing documentation, but hallucinations raise safety concerns. To mitigate these risks, we adopted a retrieval-augmented generation (RAG) approach that grounds responses in standard nursing terminologies [1]. This study aimed to develop and evaluate an RAG-based chatbot that supports nursing documentation using standard terminology and procedures.

Methods: From January 2022 to May 2024, we analyzed 70,926 unstructured nursing narratives from 76 patients with myocardial infarction who underwent percutaneous coronary intervention (PCI) at a tertiary hospital. Based on prior literature, only records from two days before to three days after PCI were included in the analysis[2,3]. The nursing records were segmented into meaningful units and used to construct a PCI nursing knowledge base grounded in NANDA–NIC–NOC (NNN), with each nursing statement linked to a corresponding SNOMED CT concept. This knowledge base served as the retrieval source for an RAG-assisted chatbot. Seventeen PCI clinical scenarios were submitted with identical prompts to three LLMs (Meta-Llama-3.1-8B, Mixtral-8x7B, and Qwen2.5-14B) to generate nursing documentation suggestions. Four nurses with at least two years of clinical experience rated each output on a 5-point Likert scale for accuracy, completeness, readability, safety, and human-centeredness[4]. Scores from the first run (run 1) were used to compare models, and the mean score across three runs was used to assess consistency.

Results: Nursing documentation most frequently focused on patient education, notification of physicians, vital sign monitoring, and confirmation of chest pain status. Across all scenarios, Qwen2.5-14B achieved the highest mean scores (accuracy: 4.69; completeness: 4.46; readability: 4.51; safety: 4.76; human-centeredness: 4.76/5), while Meta-Llama-3.1-8B and Mixtral-8x7B also scored above 4.0 in most domains. Completeness was relatively lower than the other domains for all models.

Conclusion: The RAG-based LLM chatbot using a standardized PCI nursing knowledge base generated nursing record suggestions with high accuracy and safety, primarily emphasizing patient monitoring and related interventions, which are central priorities in PCI nursing care.