OpenAI Product Pulse

Watermarking May Affect Agent Safety and Reliability

Watermarking May Affect Agent Safety and Reliability

Recent studies and external testing report that AI watermarking can change refusal behavior, tool selection, malformed tool calls, and susceptibility to prompt injection. The evidence is limited, largely involves open-weight models, and does not yet establish how OpenAI systems behave, but it raises a significant question for enterprise agents: provenance mechanisms may need dedicated behavioral and security testing.

Sources (3)
Updated Sep 17, 2026
Watermarking May Affect Agent Safety and Reliability - OpenAI Product Pulse | NBot | nbot.ai