ZICQ
中 Log in / Sign up
Newsroom Research & Papers #VLA Models #Benchmarking #Behavioral Evaluation #AI Research #arXiv

arXiv Releases ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflict

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:arXiv has released ConflictVLA-Bench, a novel benchmark designed to systematically evaluate the behavioral responses of Vision-Language-Action (VLA) models to task premise conflicts. By pairing conflict rollouts with premise-consistent reference rollouts, the benchmark assesses both outcomes and execution processes. The findings reveal that VLA models often exhibit 'Failed Persistence,' where they continue pursuing original goals despite encountering invalid premises, leading to significant drop


Background and Motivation

Vision-Language-Action (VLA) models perform strongly on manipulation tasks, but their responses to invalid task premises remain underexplored. Existing evaluations of premise conflicts often focus on terminal task outcomes, yet task failure alone cannot distinguish behavioral disengagement from continued pursuit followed by an execution error. We refer to the latter pattern as 'Failed Persistence.'

Design of ConflictVLA-Bench

To study this phenomenon, we introduce ConflictVLA-Bench, which pairs conflict rollouts with premise-consistent reference rollouts and evaluates both outcomes and execution processes. Built on LIBERO, the benchmark contains 2,826 prompt-conditioned conflict tasks spanning four conflict families, four structural configurations, and two prompt conditions.

Key Findings

  1. Reduction in Task Completion: Across all eight VLAs, invalid premises reduce original goal completion by at least 17.3 percentage points, with the reduction reaching 56.2 percentage points for OpenVLA.
  2. Failed Persistence: Even when models succeed on premise-consistent tasks and fail on their conflicting conflict tasks, they often continue to approach the original targets, retain early trajectory structure, and show limited action magnitude suppression.
  3. Insufficient Behavioral Adjustment: Explicit premise checking does not consistently produce selective and coordinated behavioral changes.

Conclusions and Implications

These findings show that terminal failure alone establishes neither behavioral disengagement nor refusal, and that outcomes alone are insufficient for VLA evaluation. ConflictVLA-Bench provides a new tool for evaluating the behavioral responses of VLA models and offers valuable data for future model improvements.

Developer Recommendations

  • Focus on Behavioral Consistency: When designing VLA models, consider the behavioral consistency when faced with conflicting premises.
  • Introduce More Complex Behavioral Evaluation Metrics: In addition to task outcomes, evaluate the behavioral process of the models.
  • Use ConflictVLA-Bench for Testing: It is recommended that developers use ConflictVLA-Bench to evaluate and improve their models' behavioral response capabilities.

Source: ArXiv AI (cs.AI) (2026-09-29)

— END —

Tags: #VLA Models #Benchmarking #Behavioral Evaluation #AI Research #arXiv

Community Comments

Loading live comments and annotations…