arXiv Releases New Research: Exploring Standalone LLMs and Multi-Step Agentic Pipelines for ICU Mortality Prediction
By Mr.Xu
Published: · 2 views
Summary:arXiv has released a new study investigating the application of Large Language Models (LLMs) in predicting ICU mortality. The research compares standalone LLMs with multi-step agentic pipelines in explaining ICU mortality predictions, finding that agentic pipelines offer advantages in enhancing clinical explanation safety and patient-specific details. However, they require attribution-based checks to ensure reliability in high-stakes risk explanations. The study, conducted on the eICU Demo datas
Background and Motivation
Machine learning models excel at predicting ICU mortality, but feature attribution methods alone often fail to provide the detailed explanations needed for clinical use. Large Language Models (LLMs) offer a potential solution, while multi-step agentic pipelines, by separating data interpretation, guideline checking, and final explanation, provide a more reliable framework for clinical applications.
Key Experiments and Results
The study, conducted on the eICU Demo dataset (2,353 ICU stays, 8.1% mortality), used the XGBoost model to achieve an AUROC of 0.855 (95% CI 0.796-0.906) and an AUPRC of 0.332 (95% CI 0.217-0.494). On a stratified 38-case explanation subset, the standalone LLM produced one explanation with explicit outcome leakage, while the four-step agentic pipeline produced none. Among the 14 cases overlapping with the SHAP review subset, the standalone LLM showed higher SHAP alignment (mean Jaccard index 0.171 vs 0.077) and direction consistency (92.9% vs 78.6%), while the agentic pipeline outperformed in guideline grounding (0.762 vs 0.143), value specificity (0.236 vs 0.143), and plausibility (0.700 vs 0.671).
Clinical Implications and Recommendations
The results suggest that agentic pipelines can enhance safety-relevant grounding and patient-specific details, but should be paired with attribution-based checks for high-stakes risk explanations. While standalone LLMs perform better in some metrics, agentic pipelines are more suitable for clinical applications due to their superior safety and interpretability.
Technical Highlights
- Multi-Step Agentic Pipeline Design: By separating data interpretation, guideline checking, and final explanation, the pipeline improves the reliability and safety of explanations.
- SHAP Alignment Analysis: The standalone LLM shows higher SHAP alignment and direction consistency, but the agentic pipeline excels in guideline grounding and value specificity.
- Clinical Applicability Assessment: The agentic pipeline is more appropriate for high-risk clinical scenarios, while the standalone LLM may be more efficient in certain cases.
Industry Impact and Future Directions
This research provides new insights into the application of LLMs in healthcare, particularly in scenarios requiring high safety and interpretability. The introduction of agentic pipelines offers a more reliable technical path for clinical decision support systems. Future research could further optimize the design of agentic pipelines and explore their potential applications in other medical fields.
— END —Source: ArXiv AI (cs.AI) (2026-08-28)
Tags: #Large Language Models #Intelligent Agents #Healthcare AI #ICU Mortality Prediction #Clinical Decision Support
Community Comments