The Mark of the Machine


Surely Anthropic did not institute its watermark to discourage people from using Claude, or to make its product less appealing—so why, then, would it choose to make its automated text more obviously traceable in the first place? The company’s official announcement explained that the change was made to comply with the E.U. A.I. Act, which requires L.L.M. providers to make A.I.-generated text identifiable as such. This identification could take the form of a watermark, metadata, or a cryptographic signal—statistical signatures that allow detection systems to determine whether text, images, or audio was created or manipulated by A.I. Google Gemini, for instance, uses a watermark called SynthID in its A.I.-generated material. OpenAI currently applies detection methods only to images and audio, but, after the E.U. A.I. Act’s “transparency obligations” went into effect, the company stated that it was “working to expand provenance measures” for ChatGPT text. These A.I. companies emphasize that tracking properties like the watermark won’t be discernible to the average reader and can’t be traced to whoever prompted the L.L.M. to generate the text—unless, of course, said person attempts to publish or present the text as his own and is outed by a detection system.

Anthropic’s motives, however, may not be limited to complying with the E.U. and upholding its stated values of safety and transparency. The company may, quite simply, want to make it easier for its models to verify human-written text and exclude A.I.-generated material when training its L.L.M.s, out of fear that training an A.I. model on A.I.-generated content might lead to a decline in reliability and quality. (The irony!) It also may be angling to assert “provenance”—a cherished word among A.I. providers—over their L.L.M.s’ output, establishing a sort of intellectual property over what their product creates, unique from its competitors. It’s hard not to perceive the whole identification operation as, in some ways, an advancement in A.I. surveillance—a way of tracking where a chatbot’s words go after they leave the machine, creating a Rolodex of use cases for future referral. “You know what this reminds me of?” the same addled Reddit poster wrote in his anti-watermark manifesto. “Those police operations that arrest the drug user and leave the dealer alone. Watermarking is the same thing.” A company like Anthropic may reasonably view watermarking as a harmless and sensible evolution: a boon to their training data, a legally compliant expansion of A.I. transparency, and a way to determine if something originated with its model. But, to its users, the stain of a watermark mainly forecasts a senseless form of social punishment, a tactic to tarnish anyone who dares use A.I.-generated material instead of plumbing the muck of their own mind.

Some critics have argued that the watermark will prove largely ineffective for flagging anything short of blatant chatbot use. When an L.L.M. like Claude generates text, it typically makes its word choices from a range of statistically probable options. With a watermark, however, the model hews toward a more predetermined set of outcomes. These outcomes can accumulate into a statistical pattern that allows a detection system to calculate the likelihood that Claude generated a piece of text. If one were to rigorously edit, or rewrite, text that the L.L.M. spits out, the detection system may struggle to pick up the pattern, but arguably that’s a feature and not a bug—isn’t catching the most obvious abuses of using A.I. for writing the point of something like a watermark, or at least a slightly more ethical internal system than the current system of nothing?

Aside from anxieties that the watermark is merely a red herring, an easy-to-evade compliance strategy, some watermark detractors fear that these statistical signatures will make chatbots worse at writing. John Gruber, a tech blogger, posited that Claude would now more readily default to watermarked words, rather than unleash the full scope of its linguistic range, which might result in lower-quality text, or in some cases unusable prose. This outcome is certainly possible if L.L.M.s begin to gravitate toward specific synonyms or phrases that hinder the explanatory prowess and precision of, say, an instruction manual or a doctor’s note—documents that are increasingly being outsourced to L.L.M.s. But many experts agree that Gruber’s concerns are overblown; Steven Murdoch, a professor of security engineering at University College London, told the Guardian that the change “probably wouldn’t have any noticeable impact.”



Source link

Posted in

Swedan Margen

I focus on highlighting the latest in business and entrepreneurship. I enjoy bringing fresh perspectives to the table and sharing stories that inspire growth and innovation.

Leave a Comment