AI Multibit Robust Watermark

Recovering content identifiers after text rewriting.

Research Intern, University of California, Santa Barbara · June 2026–present
Advisor: Prof. Yuheng Bu · Status: ARR under submission

This project develops robust multi-bit text watermarking for content attribution. Embedded identifiers can be recovered after paraphrasing without access to the original text or watermark-specific training.

The evaluated system achieved 99.83% watermark-bit injection success and 99.70% exact recovery of 32-bit identifiers in 10,000 resampled trials under the evaluated rewriting protocol.

On natural text, identifiers were recovered from all 20 evaluated Wikipedia and arXiv documents after rewriting at a target code rate of 0.2. These results support reliable identifier recovery within the evaluated setting.