ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #LLMs & Foundation Models #Context Compaction #Session Constraints #AI Evaluation #AI Governance

COMPINT: A New Benchmark for Evaluating Session Constraint Loss under Context Compaction

Avatar of Mr.Xu

By Mr.Xu

Published: · 6 views

中文阅读 (Chinese) English Version

Summary:Researchers introduce COMPINT, an evaluation suite designed to quantify the loss of Session Constraints (SCs) during context compaction in Large Language Models (LLMs). Experiments reveal that current compactors retain only 17% of injected SCs on average and often perform worse than non-compacted task execution across multi-turn chats, agentic trajectories, and long-horizon research scenarios. To address this, the team proposes an SC-aware extractor that achieves over 90% retention without modif


Background and Challenges

When processing long contexts, Large Language Model (LLM) systems often compact prior context to continue ongoing tasks. However, this compaction can lead to the loss of user-specified Session Constraints (SCs), such as 'do not delete any emails until I confirm.' This loss severely impacts the consistency of LLM behavior and, consequently, its performance in complex tasks.

COMPINT Evaluation Suite

To quantify this loss, researchers developed COMPINT, an evaluation suite designed to assess compactor performance across three long-context scenarios:

  • Multi-turn Chat: Simulates interactions between users and LLMs over multiple turns.
  • Agentic Trajectory: Evaluates LLM behavior consistency when executing complex tasks.
  • Long-horizon Research: Tests LLM performance in tasks with extended time spans.

Key Findings

  1. Significant Loss: Current compactors retain only 17% of injected SCs on average.
  2. Performance Variability: The retention rate varies significantly with different compactors, prompts, context lengths, SC phrasing, and injection locations.
  3. Systematic Defect: The loss is systematic rather than tied to any single setting.

Solution

The research team proposes an SC-aware extractor that runs alongside the compactor as a plug-and-play module. This extractor achieves over 90% retention without modifying the compactor or LLM, offering a practical solution for maintaining LLM behavior consistency.

Technical Highlights

  • Innovative Evaluation Framework: COMPINT provides the first standardized method for quantifying SC loss.
  • Plug-and-Play Module: The SC-aware extractor can be deployed without major modifications to existing systems.
  • Cross-Scenario Applicability: The method performs well across multi-turn chats, agentic trajectories, and long-horizon research scenarios.

Industry Impact and Developer Recommendations

  • Enhancing LLM Reliability: This research helps developers better understand and manage LLM behavior consistency in complex tasks.
  • Advancing AI Governance: By quantifying SC loss, researchers can more effectively formulate behavioral norms and constraints for AI systems.
  • Developer Advice: It is recommended to prioritize the integration of the SC-aware extractor in applications requiring high behavior consistency.

Conclusion

The release of COMPINT provides an important tool for evaluating and addressing SC loss during context compaction in LLMs, thereby advancing the reliability of AI systems in complex tasks.

— END —

Tags: #LLMs & Foundation Models #Context Compaction #Session Constraints #AI Evaluation #AI Governance

Community Comments

Loading live comments and annotations…