ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Medical AI #Large Language Models #Reasoning Audit #Chain-of-Thought #Perturbation Analysis

ArXiv Proposes Medical Chain-of-Thought Perturbation Audit Framework to Uncover Flaws in Medical LLM Reasoning

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:ArXiv researchers introduce a novel Medical Chain-of-Thought Perturbation Audit framework to evaluate the reliability of medical large language models (LLMs) in reasoning tasks. The method employs 30 clinically motivated perturbation operators to modify both questions and reasoning chains, paired with a joint analysis of chain updates and answer flips to classify model failure modes. Results show a Chain-Decoupling Rate (CDR) of 72.9% for clinically meaningful destructive edits, indicating that


Background and Motivation

In the medical field, Chain-of-Thought (CoT) reasoning is commonly used to demonstrate the reasoning process of models. However, current evaluation methods often treat CoT as a black box, ignoring its actual role. The ArXiv team proposes a new audit framework to systematically evaluate the reliability of medical LLMs when faced with clinically relevant modifications.

Method and Experiments

The study designs a toolkit of 30 clinically motivated perturbation operators, such as severity reversal, negation flip, demographic swap, and evidence ablation. These operators are applied to both questions and reasoning chains, and a joint analysis of chain updates and answer flips is used to classify model failure modes. Experiments were conducted on 14 LLMs across four medical QA benchmarks, yielding the following results:

  • Chain-Decoupling Rate (CDR) of 72.9%: In clinically meaningful destructive edits, most changes in the chain do not affect model outputs.
  • Chain corruption does not impact accuracy: Even when the chain is corrupted, the model's accuracy remains largely unchanged.
  • Removing CoT prompting does not reduce accuracy: Taking away the CoT prompt does not significantly affect the model's accuracy.

Additionally, two board-certified clinicians re-annotated 197 perturbed questions, and 98.5% of the original answers remained defensible.

Conclusions and Implications

The findings suggest that the chain-of-thought reasoning in current medical LLMs may be more decorative than functional. This has significant implications for the application of medical AI, suggesting that developers should focus more on the actual reasoning capabilities of models rather than relying solely on CoT as a measure of trustworthiness.

Recommendations for Developers

  • Emphasize actual reasoning capabilities: Do not rely solely on CoT as the sole indicator of model reliability.
  • Adopt stricter evaluation standards: In medical AI applications, stricter evaluation methods should be employed to ensure the model's robustness when faced with perturbations.
  • Explore new reasoning mechanisms: Consider developing new reasoning mechanisms to better support the medical decision-making process.

Source: ArXiv cs.AI (2026-08-25)

— END —

Tags: #Medical AI #Large Language Models #Reasoning Audit #Chain-of-Thought #Perturbation Analysis

Community Comments

Loading live comments and annotations…