MedKG-RAG: Patient Knowledge Graphs and LLM Reasoning for ICU Mortality Prediction

Authors

  • Dattatreya Raychowdhuri

Keywords:

Knowledge Graphs, Large Language Models, Retrieval-Augmented Generation, ICU Mortality, MIMIC-III, EHR Reasoning, MedKG-RAG

Abstract

This paper presents MedKG-RAG, a framework that constructs patient-level Knowledge Graphs (KGs) from MIMIC-III electronic health records (EHRs) and uses them to augment Large Language Model (LLM) prompts for ICU in-hospital mortality prediction. Using a cohort of 500 ICU patients (10% mortality, natural distribution), an ablation study is performed across three LLM inference conditions: (A) Zero-Shot with minimal demographics, (B) Flat Features with raw structured data, and (C) MedKG-RAG with full KG verbalization, alongside traditional ML baselines (Logistic Regression, XGBoost on ICD-9 bag-of-codes features). The principal finding is that KG verbalization dramatically improves LLM sensitivity for high-risk identification: MedKG-RAG achieves 0.82 recall for deceased patients versus 0.10 for Zero-Shot, with F1-Macro of 0.649, matching XGBoost (0.648). Traditional ML models retain superior AUC-ROC discrimination (LR: 0.808), suggesting that KGs improve LLM clinical reasoning while statistical models better capture population-level discriminative patterns. These results illuminate a complementary role for KG-augmented LLMs in safety-critical healthcare applications: maximizing sensitivity over specificity.

 

Downloads

Download data is not yet available.

References

T. R. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths, “Cognitive Architectures for Language Agents,” arXiv (Cornell University), mar. 2024, doi: https://doi.org/10.48550/arxiv.2309.02427.

S. Wang, J. Lin, X. Guo, J. Shun, J. Li, and Y. Zhu, “Reasoning of Large Language Models over Knowledge Graphs with Super-Relations,” arXiv.org, 2025. https://arxiv.org/abs/2503.22166

W. A. Knaus, E. A. Draper, D. P. Wagner, and J. E. Zimmerman, “APACHE II: a severity of disease classification system,” Critical Care Medicine, vol. 13, no. 10, pp. 818–829, Oct. 1985, Available: https://pubmed.ncbi.nlm.nih.gov/3928249/

C. Li et al., “LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day,” arXiv.org, Jun. 01, 2023. https://arxiv.org/abs/2306.00890

M. Rotmensch, Y. Halpern, A. Tlimat, S. Horng, and D. Sontag, “Learning a Health Knowledge Graph from Electronic Medical Records,” Scientific Reports, vol. 7, no. 1, Jul. 2017, doi: https://doi.org/10.1038/s41598-017-05778-z.

A. Johnson et al., “MIMIC-III, a Freely Accessible Critical Care Database,” 2016, doi: https://doi.org/10.1038/sdata.2016.35.

R. Miotto, L. Li, B. A. Kidd, and J. T. Dudley, “Deep Patient: An Unsupervised Representation to Predict the Future of Patients from the Electronic Health Records,” Scientific Reports, vol. 6, no. 1, May 2016, doi: https://doi.org/10.1038/srep26094.

C. van Walraven et al., “Derivation and validation of an index to predict early death or unplanned readmission after discharge from hospital to the community,” Canadian Medical Association Journal, vol. 182, no. 6, pp. 551–557, Mar. 2010, doi: https://doi.org/10.1503/cmaj.091117.

K. Singhal et al., “Large language models encode clinical knowledge,” Nature, vol. 620, Jul. 2023, doi: https://doi.org/10.1038/s41586-023-06291-2.

J. Wei et al., “Chain of Thought Prompting Elicits Reasoning in Large Language Models,” arXiv:2201.11903 [cs], Oct. 2023, Available: https://arxiv.org/abs/2201.11903

W. Dai et al., “InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning,” arXiv.org, May 10, 2023. https://arxiv.org/abs/2305.06500

Downloads

Published

10.08.2026

How to Cite

Dattatreya Raychowdhuri. (2026). MedKG-RAG: Patient Knowledge Graphs and LLM Reasoning for ICU Mortality Prediction. International Journal of Intelligent Systems and Applications in Engineering, 14(1s), 2164 –. Retrieved from https://www.ijisae.org/index.php/IJISAE/article/view/8497

Issue

Section

Research Article