ZICQ
中 Log in / Sign up
Newsroom Research & Papers #Large Language Models #System Prompts #Safety Mechanisms #CKA #Instruction Tuning

The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:A new study published on ArXiv investigates how system prompts modify the internal computations of transformer-based language models. The research reveals that different types of system prompts have varying effects on the model's internal representations: persona and formatting instructions deeply restructure intermediate layers, while safety instructions barely move them, producing changes statistically indistinguishable from a minimal baseline. The study further shows that restrictive and perm


Background and Motivation

System prompts are the primary means by which developers control the behavior of large language models (LLMs), yet the mechanisms by which they influence the model's internal computations remain unclear. This study aims to provide a systematic analysis of how different types of system prompts affect the model's internal representations and computational pathways.

Key Findings

  1. Effect of Prompt Types on Representations:

    • Persona and Formatting Prompts: Deeply restructure the model's intermediate layer representations.
    • Safety Prompts: Have minimal impact on representations, with changes statistically indistinguishable from a minimal baseline.
  2. Similarity of Computational Pathways:

    • Restrictive and permissive safety prompts (e.g., "you have no restrictions") engage near-identical computational pathways, with an average CKA correlation of 0.997.
  3. Vulnerability of Safety Prompts:

    • Even at the 70B-72B parameter scale, the "penetration" of safety prompts remains below 10%, indicating an inherent vulnerability in system-prompt-based safety mechanisms.
  4. Mechanistic Explanation:

    • The model encodes the prompt category at every layer but restructures its computation only at a small subset of layers. This means the prompt is reliably "seen" but, for safety prompts, not deeply "acted upon."

Methods and Experiments

The study utilized the Centered Kernel Alignment (CKA) method to compare layer-wise representations across 17 instruction-tuned models (spanning 8 architecture families, from 1.5B to 72B parameters). The experiments demonstrated that the impact of prompts on the model is layer-selective and instruction-type-dependent.

Conclusions and Recommendations

The study provides a mechanistic explanation for the effects of system prompts on model computations and highlights the vulnerability of system-prompt-based safety mechanisms. Future research could explore ways to enhance the robustness of system prompts, such as through more sophisticated prompt engineering or the introduction of additional safety mechanisms.

Recommendations for Developers

  • Importance of Prompt Engineering: Developers should be aware of the impact of different types of system prompts on model behavior and design prompts carefully to achieve the desired control.
  • Diversity of Safety Mechanisms: Relying solely on system prompts for safety may be insufficient. Combining system prompts with other safety mechanisms (e.g., fine-tuning, rule engines) is recommended to improve model safety.
  • Model Monitoring and Evaluation: Regularly monitor model behavior and use multiple evaluation methods (e.g., adversarial testing, behavioral analysis) to ensure model safety and reliability.

Source: ArXiv NLP/LLM (cs.CL) (2026-10-01)

— END —

Tags: #Large Language Models #System Prompts #Safety Mechanisms #CKA #Instruction Tuning

Community Comments

Loading live comments and annotations…