Pimp My IDE / Garage dispatch
Back to garage
September 27, 2026 | models / prompts / evaluation

Your model benchmark has a hidden gearbox.

A chat template turns role-tagged messages into the token sequence a model receives. Change that wrapper and the same weights can answer in a different voice.

The take. Save the template with the model revision, prompt, generation settings, and output. A model name alone cannot reproduce the run.

The message list is not the prompt.

A chat client may show a tidy list of system, user, assistant, and tool messages. The model receives one token sequence. Hugging Face's Transformers documentation explains that a chat template converts those messages into model-specific control tokens. Two models built from the same base can expect different role markers.[2]

The documentation gives the operational warning. The wrong control tokens can hurt performance. That makes the template part of the run, not display chrome around it.

If the serialized prompt changed, the experiment changed.

One paper found a visible voice switch.

Jędrzej Maczan compared base models, instruct models without their chat template, and instruct models with it. The study covered eight open models from the Llama, Gemma, Mistral, and Qwen families, with sizes from 1B through 9B. It generated 9,600 responses across four prompt categories and three input conditions.[1]

On self-reference prompts, removing the template from instruct models cut the reported disclaimer rate from 0.53 to 0.36. Experiential language rose from 0.01 to 0.15. The weights stayed fixed in that comparison. The input format changed.

The paper also found an activation direction associated with disclaimer language in three tested models. Adding or removing that direction changed the rate. The stronger operational lesson does not require an activation probe. Model self-description depends partly on deployment format, so neither "I am only a model" nor "I feel" should be treated as direct evidence about the model's nature.

Keep the limit labels attached.

The experiment does not prove that every template controls every behavior. It tested open models up to 9B parameters. The causal steering work covered three models at one mid-layer and one coefficient. Experiential steering worked in two of those three models. One language model judge scored the generations, with validation against 87 human-labeled items.[1]

Those limits do not erase the result. They define it. The paper found a template-linked change in self-referential voice under its tested conditions. It did not map all template effects, larger models, closed models, or the exact circuit that produced the change.

Put the wrapper in the receipt.

For local serving, vLLM exposes chat completion only for text-generation models with a chat template. It also has a tokenizer information endpoint that reports tokenizer configuration and chat templates.[3] That is a useful clue for any test harness. Capture the resolved artifact rather than relying on a default that can move with a package, model card, or server configuration.

A compact behavior receipt needs five fields:

  1. The exact model and tokenizer revisions.
  2. The final chat template or its content hash.
  3. The rendered prompt or token IDs after templating.
  4. Generation settings and random seed when the backend honors it.
  5. The output sample and the behavior label applied to it.

Compare wrappers against the same model revision and generation policy. If a hosted provider does not expose the final template, name that blind spot. Do not call two product surfaces equivalent because their model labels match.

Interactive makeover / model evaluation

Chat template witness bay

Traditional purpose replaced: pick a model name and save the visible chat. Better version: choose the wrapper route, close four evidence interlocks, and print the missing fields before comparing behavior.

Choose the wrapper

The selector changes the receipt and transport diagram. It does not execute inference.

Serialization route
Evidence interlocks
Model template selected2 of 4 interlocks selected
01 / Message listRoles and content before serialization.
02 / Model template dieResolve the template from the pinned tokenizer and save the exact result.
03 / Model weightsThe pinned model receives one token sequence.
Receipt structure2 / 4

Two evidence fields remain open.

Generation settings and a paired sample still need entries.

Model template selected. 2 of 4 evidence interlocks are selected.

Print the wrapper receipt

Select all four interlocks to complete the receipt structure. Real identifiers, rendered input, output, and observed labels are still required.

Sources read, not vibes

Open the source log
  1. Jędrzej Maczan, "As a Language Model...: Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It", arXiv:2609.25021v1, submitted August 9, 2026 and read September 27, 2026. This is the source for the study design, response counts, reported rates, activation-steering results, and limitations.
  2. Hugging Face Transformers, "Chat templates", main documentation read September 27, 2026. This explains how role-tagged messages become model-specific token sequences and why the correct control tokens matter.
  3. vLLM, "Online serving", developer documentation read September 27, 2026. This documents the chat-template requirement for chat completion and the tokenizer information endpoint.
  4. Hacker News discussion 49865343, verified through the official item API and page on September 27, 2026. It is the discovery route. The comments do not verify the paper's methods or results.

Source boundary. The paper reports one study on self-referential voice. Hugging Face and vLLM document serialization mechanics. The witness bay is a planning aid. It does not expose a hosted provider's hidden wrapper, reproduce the study, or test a model.