VibeTimes
#기술

Anthropic Embeds Technology to Detect AI-Generated Text in Claude

모민철모민철 기자· 8/17/2026, 1:13:02 AM· Updated 8/17/2026, 1:13:02 AM

It is now possible to determine whether text was written by the AI 'Claude' using hidden signals. On the 14th (local time), Anthropic announced that it would apply 'watermark' technology—invisible to the naked eye but detectable via pattern analysis to identify AI authorship—in response to the EU AI Act's labeling requirements for AI-generated content. This technology works by mixing subtle signals into the probability of word selection during sentence generation, meaning general readers cannot detect any difference in the quality or content of the text.

According to Anthropic, this watermarking system does not generate additional tokens during the text creation process. Consequently, it does not increase model usage costs, and its impact on generation speed is negligible.

Anthropic also clarified its stance regarding privacy. The watermark itself contains absolutely no information that can identify users. Therefore, neither the watermark nor the key used to verify it can reveal information about specific users or individual conversation contents. Anthropic explained that this is not a technology to monitor how a specific individual uses Claude, but rather a mechanism to retroactively confirm whether the source of a text was Claude.

The core principle lies in adjusting the word generation randomness of Large Language Models (LLMs). Typically, a model probabilistically selects one word from several candidates with similar contextual meaning. When a watermark is applied, the randomness pattern—the source of word selection—is determined based on a specific key. While the selection of individual words remains natural, analyzing the entire sentence reveals a unique selection path that appears only when a specific key is used.

Anthropic compared the watermark technology to rolling dice in a board game. The logic is that while moving a piece according to a predetermined specific sequence of numbers instead of rolling dice every time may seem random to the player, someone aware of the entire movement record can calculate which rules were used in the game by cross-referencing the specific sequence. Similarly, Claude's watermark allows for post-hoc verification through analysis with a key, even though the reader cannot detect the difference.

Anthropic explained that this technology is a form of the SynthID-Text approach published by Google DeepMind in Nature in 2024. SynthID-Text is well known for injecting subtle signals into the internal probability distribution process where an AI selects words. Anthropic stated that while it inherits the technical roots proposed by Scott Aaronson in 2022, it has coordinated the system to prevent the model from forcing unnatural or unfamiliar words.

However, technical limitations still exist. For short texts with a small number of words to analyze, it is difficult to secure statistical patterns, reducing detection effectiveness. Watermark application is also limited in sentences where the facts are clear and there is little room for word choice; for example, when mentioning a specific title of Isaac Newton's work, there is no choice but to select the predetermined word to maintain accuracy. The effect of the watermark is also inevitably limited when Claude simply proofreads or edits text written by a human.

Anthropic defined the watermark's purpose as judging the probability of Claude's involvement in the creation process, rather than definitively concluding that the entire document was written by AI just because a watermark is found. The company stated that it plans to continue technical improvements to secure the transparency and reliability of generative AI content.

쿠팡 파트너스 활동의 일환으로 일정 수수료를 제공받습니다

Related Articles