ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Autonomous Driving #Semantic Understanding #Knowledge Graph #Vision-Language Model #Reasoning

arXiv Introduces Deterministic Predicate Framework for Explainable Driving Scene Reasoning

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:A new research paper on arXiv introduces a deterministic multi-dataset predicate framework for driving-scene understanding. This framework derives semantic relations from measurable geometric, kinematic, temporal, map, and traffic-control evidence and materializes them in a common Predicate Knowledge Graph (Predicate KG). Experimental results show that the framework achieves macro F1 scores of 0.94 on nuPlan and 0.93 on nuScenes, demonstrating significant improvements in semantic consistency and


Background and Motivation

In recent years, Vision-Language Models (VLMs) have seen increasing use in driving-scene understanding. However, the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation, which raises concerns about their interpretability and reliability. To address this challenge, researchers have introduced a deterministic multi-dataset predicate framework that generates semantic relations from measurable evidence and materializes them in a common Predicate Knowledge Graph (Predicate KG).

Technical Highlights

  1. Multi-Dataset Interface Design: The framework uses dataset-specific interfaces to recover the required scene information, while predicate definitions remain unchanged across datasets like nuPlan and nuScenes, ensuring semantic consistency.
  2. Predicate Knowledge Graph (Predicate KG): By constructing a unified Predicate KG, the researchers provide a systematic representation and validation of driving-scene semantics.
  3. Quantitative Semantic Validation: The framework achieves macro F1 scores of 0.94 on nuPlan and 0.93 on nuScenes across 200 scenarios, demonstrating its superiority in semantic consistency.
  4. LLaVA-OneVision-7B Model Evaluation: In nine NuPlanQA subtasks, Predicate KG outperforms in seven, including Traffic Light (from 53.2% to 71.5%), Situation Assessment (from 76.2% to 86.1%), and Action Recommendation (from 82.9% to 89.0%).

Industry Impact

This research offers a new technical pathway for semantic understanding and reasoning in autonomous driving systems. By reducing reliance on visual input, Predicate KG enhances model performance in complex scenarios and improves interpretability. This advancement holds significant implications for the development of autonomous vehicles, intelligent transportation systems, and related fields.

Recommendations for Developers

  • Integrate Predicate KG: Developers can integrate Predicate KG into existing autonomous driving systems to improve semantic understanding and reasoning capabilities.
  • Explore Multi-Dataset Interface Design: Researchers can draw inspiration from the framework's design to ensure semantic consistency and scalability in multi-dataset interfaces.
  • Test with LLaVA-OneVision-7B: Developers can further test and validate Predicate KG using the LLaVA-OneVision-7B model to explore its performance in different scenarios.

Source: ArXiv Machine Learning (cs.LG) (2026-09-29)

— END —

Tags: #Autonomous Driving #Semantic Understanding #Knowledge Graph #Vision-Language Model #Reasoning

Community Comments

Loading live comments and annotations…