OpenAI published a technical report on 5 October 2026 on "textGrain", a watermark for AI text. The company will add it automatically to ChatGPT and Codex text in the European Union, with the launch planned "in the coming weeks". The mark is a statistical pattern in word choice, not a hidden symbol or an invisible space, so it cannot prove anything for certain about who wrote a text.
The move follows Article 50 of the EU Artificial Intelligence Act. From 2 August 2026, it requires systems that generate text, images, audio and video to label their output in a machine-readable way, so it can be recognised as made by AI. On 10 June 2026 the European Commission published the final Code of Practice on these issues. The code is voluntary, but companies can use it to show they follow the rule. It says no single technique meets the requirements alone, so providers must use at least two layers of marking, for example metadata and a watermark.
A statistical pattern in word choice
A language model writes word by word, and each time it picks from several fitting options. "textGrain" replaces part of this randomness with values worked out from a secret key and the words before. Anyone with the key can check whether a text follows this pattern more often than chance would explain. Copying and pasting an unchanged answer does not remove the mark.
The text looks the same, so the mark cannot be seen in a text editor or an email. OpenAI says the signal is too weak in short answers of under about 150 English words. It also says that in code, maths and strictly factual answers, the model has little room to reword, so the mark is hard to add there.
The detector works best on long texts. On passages of 400 tokens (the word fragments a model works with), detection is about 95%. At 200 tokens it is about 80%. It also depends on language, ranging from 42.2% for Romanian to 69.0% for Spanish. The sources give no figures for Bulgarian.
Editing erases the mark
If 25% of the words are replaced, detection falls to 17%. OpenAI admits that rewording or translation erases the mark. A study of SynthID, Google's similar technology, found the same: accuracy drops sharply after light rewording and translation. In another attack, where AI text is mixed with a lot of human text, more than half of the unmarked texts were wrongly flagged as marked.
A pupil or student who rewrites an answer in their own words will most likely leave no trace. A missing mark does not prove a text is human. A text in which ChatGPT only fixed the spelling may carry a mark. Finding one shows that AI took part, but not how much the human edited.
Detector open to approved organisations
The detector can be wrong both ways: it may find a mark where there is none, or miss one that is there. It is not public for now. OpenAI gives it only to approved researchers and expert organisations, including Cornell University, ETH Zurich and KInIT. The code says regulators, law enforcement bodies, media and fact-checkers must get access later.
The EU rule also says anyone who publishes AI text on matters of public interest must label it. The exception is text that has been through human editing, where someone has taken editorial responsibility.
Anthropic got there first. On 11 August 2026 it said it was putting an invisible mark in text from its models released since 2 August 2026. It warned that the mark does not prove authorship, and that heavy editing, rewording and short passages erase it. Critics say the EU rule requires the mark to be hard to separate from the content, while text marks are easy to remove. Access to the programming interface for customers worldwide stays optional and is off by default. For systems already running, the deadline to comply is 2 December 2026.
Comments (1)
Ей, ама хайде сега! "Статистически знак"?! Сериозно ли? Някой да обясни НА баба ми к'во е тоя статистически знак - все едно ще го татуират на текста. И какво печелим от това всъщност? Че ако някой си е написал нещо и после го прегледа, бе, изчезва знака! Каква работа свършваме тогава