The development of Large Language Models (LLMs) has achieved remarkable progress, with their language understanding capabilities increasingly approaching human-like performance and being widely applied. However, with the enhancement of model capabilities, issues such as the misuse of LLMs and copyright infringement have become more pronounced, leading to multiple challenges in social, legal, and technical domains. As a result, attention has been focused on model watermarking techniques, which embed imperceptible yet machine-detectable digital watermarks in model outputs, enabling content traceability. In this paper, we propose a water-marking method specifically designed for LLM-generated text, embedding watermark information through the adjustment of the Window Average Space Value (WASV) within the text. This method allows for multi-bit watermark embedding, ensuring the watermark’s invisibility by fine-tuning word replacement while maintaining low perplexity. Additionally, a lightweight watermark detection method is achieved using a sliding window and dynamic programming matching technique for watermark extraction. Our experiments on various LLMs demonstrate that our approach achieves perplexity comparable to single-bit watermark schemes. Moreover, the watermark-generated text exhibits strong robustness while remaining easily detectable.
Paper
Full text
A Multi-bit Robust LLM Watermark based on Window Average Space Value
Semantic Scholar · 2025
Abstract
The development of Large Language Models (LLMs) has achieved remarkable progress, with their language understanding capabilities increasingly approaching human-like performance and being widely applied. However, with the enhancement of model capabilities, issues such as the misuse of LLMs and copyright infringement have become more pronounced, leading to multiple challenges in social, legal, and technical domains. As a result, attention has been focused on model watermarking techniques, which embed imperceptible yet machine-detectable digital watermarks in model outputs, enabling content traceability. In this paper, we propose a water-marking method specifically designed for LLM-generated text, embedding watermark information through the adjustment of the Window Average Space Value (WASV) within the text. This method allows for multi-bit watermark embedding, ensuring the watermark’s invisibility by fine-tuning word replacement while maintaining low perplexity. Additionally, a lightweight watermark detection method is achieved using a sliding window and dynamic programming matching technique for watermark extraction. Our experiments on various LLMs demonstrate that our approach achieves perplexity comparable to single-bit watermark schemes. Moreover, the watermark-generated text exhibits strong robustness while remaining easily detectable.