Latest from MIT Tech Review – It’s easy to tamper with watermarks from AI-generated text

Watermarks for AI-generated text are easy to remove and can be stolen and copied, rendering them useless, researchers have found. They say these kinds of attacks discredit watermarks and can fool people into trusting text they shouldn’t.

Watermarking works by inserting hidden patterns in AI-generated text, which allow computers to detect that the text comes from an AI system. They’re a fairly new invention, but they have already become a popular solution for fighting AI-generated misinformation and plagiarism. For example, the European Union’s AI Act, which enters into force in May, will require developers to watermark AI-generated content. But the new research shows that the cutting edge of watermarking technology doesn’t live up to regulators’ requirements, says Robin Staab, a PhD student at ETH Zürich, who was part of the team that developed the attacks. The research is yet to be peer reviewed.

AI language models work by predicting the next likely word in a sentence, generating one word at a time on the basis of those predictions. Watermarking algorithms for text divide the language model’s vocabulary into words on a “green list” and a “red list,” and then make the AI model choose words from the green list. The more words in a sentence that are from the green list, the more likely it is that the text was generated by a computer. Humans tend to write sentences that include a more random mix of words.

The researchers tampered with five different watermarks that work in this way. They were able to reverse-engineer the watermarks by using an API to access the AI model with the watermark applied and prompting it many times, says Staab. The responses allow the attacker to “steal” the watermark by building an approximate model of the watermarking rules. They do this by analyzing the AI outputs and comparing them with normal text.

Once they have an approximate idea of what the watermarked words might be, this allows the researchers to execute two kinds of attacks. The first one, called a spoofing attack, allows malicious actors to use the information they learned from stealing the watermark to produce text that can be passed off as being watermarked. The second attack allows hackers to scrub AI-generated text from its watermark, so the text can be passed off as human-written.

The team had a roughly 80% success rate in spoofing watermarks, and an 85% success rate in stripping AI-generated text of its watermark.

Researchers not affiliated with the ETH Zürich team, such as Soheil Feizi, an associate professor and director of the Reliable AI Lab at the University of Maryland, have also found watermarks to be unreliable and vulnerable to spoofing attacks.

The findings from ETH Zürich confirm that these issues with watermarks persist and extend to the most advanced types of chatbots and large language models being used today, says Feizi.

The research “underscores the importance of exercising caution when deploying such detection mechanisms on a large scale,” he says.

Despite the findings, watermarks remain the most promising way to detect AI-generated content, says Nikola Jovanović, a PhD student at ETH Zürich who worked on the research.

But more research is needed to make watermarks ready for deployment on a large scale, he adds. Until then, we should manage our expectations of how reliable and useful these tools are. “If it’s better than nothing, it is still useful,” he says.

Latest from MIT : Building an understanding of how drivers interact with emerging vehicle technologies

As the global conversation around assisted and automated vehicles (AVs) evolves, the MIT Advanced Vehicle Technology (AVT) Consortium continues to lead cutting-edge research aimed at understanding how drivers interact with emerging vehicle technologies. Since its launch in 2015, the AVT Consortium — a global academic-industry collaboration on developing a data-driven understanding of how drivers respond…

Artificial Intelligence

Latest from MIT : LLMs develop their own understanding of reality as their language abilities improve

Ask a large language model (LLM) like GPT-4 to smell a rain-soaked campsite, and it’ll politely decline. Ask the same system to describe that scent to you, and it’ll wax poetic about “an air thick with anticipation” and “a scent that is both fresh and earthy,” despite having neither prior experience with rain nor a…

Artificial Intelligence

Latest from MIT : A creation story told through immersive technology

In the beginning, as one version of the Haudenosaunee creation story has it, there was only water and sky. According to oral tradition, when the Sky Woman became pregnant, she dropped through a hole in the clouds. While many animals guided her descent as she fell, she eventually found a place on the turtle’s back….

Artificial Intelligence

Latest from MIT Tech Review – Can AI help DOGE slash government budgets? It’s complex.

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. No tech leader before has played the role in a new presidential administration that Elon Musk is playing now. Under his leadership, DOGE has entered offices in a half-dozen agencies and counting,…

Artificial Intelligence

Latest from MIT : Pattie Maes receives ACM SIGCHI Lifetime Research Award

Pattie Maes, the Germeshausen Professor of Media Arts and Sciences at MIT and head of the Fluid Interfaces research group within the MIT Media Lab, has been awarded the 2025 ACM SIGCHI Lifetime Research Award. She will accept the award at CHI 2025 in Yokohama, Japan this April. The Lifetime Research Award is given to individuals whose research…

Artificial Intelligence

Latest from MIT Tech Review – How do you teach an AI model to give therapy?

On March 27, the results of the first clinical trial for a generative AI therapy bot were published, and they showed that people in the trial who had depression or anxiety or were at risk for eating disorders benefited from chatting with the bot. I was surprised by those results, which you can read about…

Similar Posts