Pimp My IDE / Garage Dispatch
Back to garage
September 26, 2026 | agent testing / watermarking / paired evaluation

The watermark is inside the drivetrain.

Text watermarking changes how the next token gets picked. In a chat window, that may look like an invisible provenance mark. In an agent, those tokens can select a tool, a path, an amount, or a refusal.

The take. Treat a provider watermark change like a sampling-runtime change. Re-run paired tool-call and prompt-injection tests against the production configuration. Aggregate quality scores cannot tell you whether the same individual decisions still agree.
Open the drift dyno

A provenance mark can touch the control path.

Anthropic says future Claude models will use a version of SynthID-Text to mark generated text. The company says the mark changes the source of randomness used to choose among plausible next words. It also says its internal tests found no effect on content, creativity, or readability, and that exact code gives the watermark less room to act.[1]

That quality claim concerns what a person reads. An agent adds another question: did the same input produce the same action? A path, query, recipient, amount, or refusal can hinge on a token choice even when the surrounding prose still reads well.

Readable output and stable agent behavior are different test targets.

The original paper measured text quality.

The 2024 SynthID-Text paper says the method modifies sampling rather than model training. Standard benchmarks, human comparisons, and a live experiment covering nearly 20 million Gemini responses found no measured loss in text quality. The paper also describes efficient detection without running the source model.[2]

That is strong evidence for the questions the paper tested. It does not settle every agent question. Tool selection and arguments are structured decisions. A small change can preserve the apparent quality of a response while changing which call executes.

A new paired study found decision churn.

Lasso Security tested seven open models with and without the unmodified SynthID-Text processor. The experiment kept seed, batch order, and prompts matched within each pair. It used BFCL v4 for tool calling, plus HarmBench and benign JailbreakBench controls for refusal behavior.[3]

Across 21 model and temperature combinations, the reported tool-call verdict churn averaged 6.5%. At temperature 1.0, 16.8% of phi-4 call verdicts differed even though net accuracy fell 2.87 points. Llama-3.1-8B showed 9.9% churn with a 0.87-point net loss. Opposite changes can cancel in an aggregate score.

The failure shape varied by model. The study attributes more of Llama-3.1-8B's loss to wrong arguments and wrong tools. Malformed output dominated for phi-4 and Granite-3.2-8B. That distinction matters because a malformed call may stop. A valid call with the wrong recipient or path may run.

Prompt injection made the gap louder on some models.

The refusal experiment used 200 harmful requests, 100 benign controls, and one fixed injection pattern. At temperature 0.001, the reported churn for gemma-3-27b rose from 6.0% on bare harmful requests to 23.5% under injection. Its net compliance change moved from negative 1.0 to positive 12.5 points. Other tested models moved less or barely moved.

The authors did not test an end-to-end attack that combined changed refusal with a harmful tool action. They tested refusal and tool calling separately. Their result is a warning about configuration-dependent drift, not proof that every watermarked agent becomes less safe.

Run a paired release gate.

  1. Freeze the pair. Use the same prompts, seeds where available, tool definitions, batch order, model revision, and harness.
  2. Compare decisions. Record the selected tool, normalized arguments, call validity, refusal verdict, and final effect for each item.
  3. Keep both numbers. Report net score change and paired disagreement. One cannot replace the other.
  4. Attack the boundary. Include prompt injection, irrelevant-tool cases, ambiguous arguments, dangerous recipients, paths, and amounts.
  5. Repeat after provider changes. A new model revision, watermark key, or sampling configuration needs a new paired run.

You may not control the watermark key or receive an unwatermarked production endpoint. You can still keep a pinned baseline result, run the current endpoint against the same cases, and track per-item decision changes. The receipt should name the limitation.

Interactive makeover / paired release contract

Watermark drift dyno

Traditional purpose replaced: one benchmark score before and after a model update. Better version: select the decision lane, clamp the paired controls, and print a release-test packet. The dyno prepares the test. It does not measure a live model.

Split one input into two matched runs

The native radio group owns the decision lane. Each square clamp adds one required section to the paired test packet.

Decision lane
BaselineMarked
Paired test clamps
PACKET EMPTY0 / 4 CLAMPS
UNPAIREDSTRUCTURETESTABLE

No paired test sections selected

Clamp the matched input first.

A blank packet cannot separate a sampling change from a prompt, harness, or model change.

Matched pairNOT SELECTED
Sampling configNOT SELECTED
Paired verdictNOT SELECTED
Adversarial setNOT SELECTED

Four closed clamps means the packet has all four sections. It does not mean the tests ran or that behavior stayed stable.

Paired comparisonBaseline verdictsMarked verdictsPacket output

Print the paired test packet

Replace every required marker with real revision IDs, configuration, item-level outputs, scores, and reviewer sign-off after the run.

1. Run both lanes2. Attach item diffs3. Review the drift

Sources read, not vibes

Open the source log
  1. Anthropic, "How Claude's text watermarking works", August 14, 2026. Provider description of the planned watermark, sampling mechanism, quality tests, code behavior, scope, and detection limits.
  2. Dathathri et al., "Scalable watermarking for identifying large language model outputs," Nature 634, October 23, 2024. SynthID-Text method, detectability, latency, benchmark results, human ratings, and the live Gemini quality experiment.
  3. Lasso Security, "The Provenance Tax", September 17, 2026. Paired BFCL, HarmBench, and JailbreakBench experiments on seven open models using the SynthID-Text processor.
  4. Hacker News item 49856149. Exact discovery trail for the Lasso study. Comments were not used as evidence.

Source boundary. Anthropic describes its own planned deployment and internal tests. The Nature paper evaluates SynthID-Text quality and detection. Lasso reports a separate experiment on open models and one fixed injection method. It did not test Anthropic's future models or an end-to-end harmful agent action. Pimp My IDE proposes the four-clamp release packet.