Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents
Suxin Ji, Hungtao Wan, Shaoxuan Chen, An Zhang
- Published
- Oct 6, 2026 — 16:50 UTC
Problem
Prior watermarking techniques for agent models are susceptible to desynchronization and forgery, which undermines their reliability in ensuring provenance. This paper presents a solution to these vulnerabilities through a novel watermarking approach. The work is a preprint and has not undergone peer review.
Method
The authors propose Semantic Behavioral Watermarking (SBW), which utilizes a watermarking algorithm based on semantic action clusters conditioned on historical actions. Key features of the method include:
- Keyed Binning: This employs collision-resistant binning with unpredictable fresh-bucket assignments, operating within the random-oracle model to enhance security against forgery.
- Data Utilization: The method is evaluated on five agent models ranging from 3B to 14B parameters from four different vendors. The datasets used include ToolBench, which consists of 600 trajectories per model, and ALFWorld, with 100 episodes per model.
- Training Compute and Loss Function: Specific details regarding training compute and the loss function used are not disclosed in the paper.
Results
The performance of the proposed SBW method is evaluated against various benchmarks:
- ToolBench Detection (Rewriting): Achieved a detection rate of 0.49-0.66 for cluster-level watermarking compared to 0.05-0.17 for exact-symbol detection at a 1% false positive rate (FPR).
- ToolBench Choice Agreement: Demonstrated a choice agreement of 72-83% versus 22-27% for logit biasing.
- ALFWorld Detection: Achieved a detection rate of 0.92-0.97 compared to 0.00-0.01 for the baseline.
- Keyed Binning Forgery Reduction: Successfully reduced forgery rates from 100% to a false-positive floor at the primary operating point (bge, r=64).
- Chained Replay Detection: Achieved detection rates of 0.76-0.98 across the five models tested.
- Paraphrase Robustness Cost: The cost of maintaining paraphrase robustness is approximately half of the per-step watermark capacity.
Limitations
The authors acknowledge that while the chained replay detection remains effective (0.76-0.98), it is still detectable, indicating a potential area for improvement. Additionally, the requirement for paraphrase robustness results in a reduction of the watermark capacity, which may limit its effectiveness in certain applications.
Why it matters
The implications of this work are significant for the development of secure and reliable LLM agents. By addressing the vulnerabilities of previous watermarking methods, SBW provides a more robust framework for ensuring provenance, which is critical for trust and accountability in AI systems. This advancement could pave the way for further research into watermarking techniques that balance robustness and capacity, enhancing the integrity of AI-generated content.
By Turing Wire Research Desk · Oct 6, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
