ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #TempJail #I2V Models #Security Research #ArXiv #Temporal Jailbreak

ArXiv Introduces TempJail: Temporal Jailbreak Attacks on Image-to-Video Generation Models

Avatar of Mr.Xu

By Mr.Xu

Published: · 4 views

中文阅读 (Chinese) English Version

Summary:The ArXiv team introduces TempJail, a novel framework for temporal jailbreak attacks on image-to-video (I2V) generation models. Unlike traditional studies that focus on single-frame violations, TempJail uncovers a temporal vulnerability in I2V systems, where unsafe semantics may emerge from semantic compositions over time. The framework decomposes malicious captions into initial frame visual conditions and temporal text instructions, and employs techniques for semantic camouflage on both image a


Key Breakthroughs

The ArXiv team introduces TempJail, a novel framework for temporal jailbreak attacks on image-to-video (I2V) generation models. The key breakthroughs include:

  • Uncovering Temporal Vulnerabilities: Unlike traditional studies that focus on single-frame violations, TempJail reveals a temporal vulnerability in I2V systems, where unsafe semantics may emerge from semantic compositions over time.
  • Innovative Attack Framework: The framework decomposes malicious captions into initial frame visual conditions and temporal text instructions, and employs techniques for semantic camouflage on both image and text sides to achieve more effective attacks.
  • Experimental Validation: Experiments show that TempJail improves attack success rates by 23.3% under GPT-5.2 evaluation and 22.0% under human evaluation compared to state-of-the-art methods.

Technical Highlights

  1. Temporal Abstraction: The framework decomposes target malicious descriptions into initial frame visual conditions and temporal text instructions to effectively manipulate the temporal dimension.
  2. Semantic Camouflage: On the image side, semantic injection is modeled as controlled latent perturbation in diffusion sampling, while on the text side, captions are rewritten into innocuous 'subject-action-scene' templates to bypass safety filters.
  3. Black-Box Inference: In the black-box inference phase, the two modalities work together to gradually trigger malicious semantics over time.

Industry Impact

  • Safety Evaluation: TempJail provides a new perspective for the safety evaluation of I2V models, revealing potential risks in the temporal dimension of existing models.
  • Model Improvement Recommendations: The study suggests that I2V model developers should strengthen temporal safety measures, such as introducing time-aware filtering mechanisms.
  • Future Research Directions: Future research can further explore temporal jailbreak attacks on other modalities (e.g., audio) and develop more robust defense mechanisms.

Developer Recommendations

  • Strengthen Temporal Safety Measures: Developers should consider incorporating time-aware filtering mechanisms into their models to detect and prevent temporal jailbreak attacks.
  • Multimodal Safety Evaluation: It is recommended to conduct multimodal safety evaluations of I2V models, including comprehensive assessments of audio, text, and image modalities.
  • Continuous Monitoring and Updates: Regularly monitor the security performance of models and update and optimize them based on the latest research findings.

Source: ArXiv cs.CV (2026-08-27)

— END —

Tags: #TempJail #I2V Models #Security Research #ArXiv #Temporal Jailbreak

Community Comments

Loading live comments and annotations…