The message list is not the prompt.
A chat client may show a tidy list of system, user, assistant, and tool messages. The model receives one token sequence. Hugging Face's Transformers documentation explains that a chat template converts those messages into model-specific control tokens. Two models built from the same base can expect different role markers.[2]
The documentation gives the operational warning. The wrong control tokens can hurt performance. That makes the template part of the run, not display chrome around it.
If the serialized prompt changed, the experiment changed.
One paper found a visible voice switch.
Jędrzej Maczan compared base models, instruct models without their chat template, and instruct models with it. The study covered eight open models from the Llama, Gemma, Mistral, and Qwen families, with sizes from 1B through 9B. It generated 9,600 responses across four prompt categories and three input conditions.[1]
On self-reference prompts, removing the template from instruct models cut the reported disclaimer rate from 0.53 to 0.36. Experiential language rose from 0.01 to 0.15. The weights stayed fixed in that comparison. The input format changed.
The paper also found an activation direction associated with disclaimer language in three tested models. Adding or removing that direction changed the rate. The stronger operational lesson does not require an activation probe. Model self-description depends partly on deployment format, so neither "I am only a model" nor "I feel" should be treated as direct evidence about the model's nature.
Keep the limit labels attached.
The experiment does not prove that every template controls every behavior. It tested open models up to 9B parameters. The causal steering work covered three models at one mid-layer and one coefficient. Experiential steering worked in two of those three models. One language model judge scored the generations, with validation against 87 human-labeled items.[1]
Those limits do not erase the result. They define it. The paper found a template-linked change in self-referential voice under its tested conditions. It did not map all template effects, larger models, closed models, or the exact circuit that produced the change.
Put the wrapper in the receipt.
For local serving, vLLM exposes chat completion only for text-generation models with a chat template. It also has a tokenizer information endpoint that reports tokenizer configuration and chat templates.[3] That is a useful clue for any test harness. Capture the resolved artifact rather than relying on a default that can move with a package, model card, or server configuration.
A compact behavior receipt needs five fields:
- The exact model and tokenizer revisions.
- The final chat template or its content hash.
- The rendered prompt or token IDs after templating.
- Generation settings and random seed when the backend honors it.
- The output sample and the behavior label applied to it.
Compare wrappers against the same model revision and generation policy. If a hosted provider does not expose the final template, name that blind spot. Do not call two product surfaces equivalent because their model labels match.