A provenance mark can touch the control path.
Anthropic says future Claude models will use a version of SynthID-Text to mark generated text. The company says the mark changes the source of randomness used to choose among plausible next words. It also says its internal tests found no effect on content, creativity, or readability, and that exact code gives the watermark less room to act.[1]
That quality claim concerns what a person reads. An agent adds another question: did the same input produce the same action? A path, query, recipient, amount, or refusal can hinge on a token choice even when the surrounding prose still reads well.
Readable output and stable agent behavior are different test targets.
The original paper measured text quality.
The 2024 SynthID-Text paper says the method modifies sampling rather than model training. Standard benchmarks, human comparisons, and a live experiment covering nearly 20 million Gemini responses found no measured loss in text quality. The paper also describes efficient detection without running the source model.[2]
That is strong evidence for the questions the paper tested. It does not settle every agent question. Tool selection and arguments are structured decisions. A small change can preserve the apparent quality of a response while changing which call executes.
A new paired study found decision churn.
Lasso Security tested seven open models with and without the unmodified SynthID-Text processor. The experiment kept seed, batch order, and prompts matched within each pair. It used BFCL v4 for tool calling, plus HarmBench and benign JailbreakBench controls for refusal behavior.[3]
Across 21 model and temperature combinations, the reported tool-call verdict churn averaged 6.5%. At temperature 1.0, 16.8% of phi-4 call verdicts differed even though net accuracy fell 2.87 points. Llama-3.1-8B showed 9.9% churn with a 0.87-point net loss. Opposite changes can cancel in an aggregate score.
The failure shape varied by model. The study attributes more of Llama-3.1-8B's loss to wrong arguments and wrong tools. Malformed output dominated for phi-4 and Granite-3.2-8B. That distinction matters because a malformed call may stop. A valid call with the wrong recipient or path may run.
Prompt injection made the gap louder on some models.
The refusal experiment used 200 harmful requests, 100 benign controls, and one fixed injection pattern. At temperature 0.001, the reported churn for gemma-3-27b rose from 6.0% on bare harmful requests to 23.5% under injection. Its net compliance change moved from negative 1.0 to positive 12.5 points. Other tested models moved less or barely moved.
The authors did not test an end-to-end attack that combined changed refusal with a harmful tool action. They tested refusal and tool calling separately. Their result is a warning about configuration-dependent drift, not proof that every watermarked agent becomes less safe.
Run a paired release gate.
- Freeze the pair. Use the same prompts, seeds where available, tool definitions, batch order, model revision, and harness.
- Compare decisions. Record the selected tool, normalized arguments, call validity, refusal verdict, and final effect for each item.
- Keep both numbers. Report net score change and paired disagreement. One cannot replace the other.
- Attack the boundary. Include prompt injection, irrelevant-tool cases, ambiguous arguments, dangerous recipients, paths, and amounts.
- Repeat after provider changes. A new model revision, watermark key, or sampling configuration needs a new paired run.
You may not control the watermark key or receive an unwatermarked production endpoint. You can still keep a pinned baseline result, run the current endpoint against the same cases, and track per-item decision changes. The receipt should name the limitation.