El watermark de Claude y por qué la trazabilidad no puede vivir en el outputClaude's Watermark and Why Traceability Can't Live in the Output
Analista de Transfer Pricing. Investigador en gobernanza algorítmica aplicada al cumplimiento fiscal internacional. Creador del Sistema TRIA.Transfer Pricing Analyst. Researcher in algorithmic governance applied to international tax compliance. Creator of Sistema TRIA.
Anthropic empezó a incorporar watermarks invisibles en el texto que genera Claude, en cumplimiento del Código de Prácticas de la UE bajo el AI Act. El mecanismo es elegante: en cada punto donde el modelo elige entre palabras igualmente válidas —"nublado" o "gris" después de "el clima estaba frío y..."—, esa elección deja de ser aleatoria y pasa a depender de una clave, de modo que quien tiene esa clave puede verificar si un texto fue producido por Claude. No se inserta nada, no hay caracteres ocultos, no cambia el costo ni la velocidad.Anthropic began embedding invisible watermarks in text generated by Claude, in compliance with the EU's Code of Practice under the AI Act. The mechanism is elegant: at every point where the model chooses between equally valid words — "overcast" or "grey" after "the weather today was cold and..." — that choice stops being random and instead depends on a key, so that whoever holds the key can verify whether a text was produced by Claude. Nothing gets inserted, there are no hidden characters, and it doesn't change cost or speed.
Es, técnicamente, un logro fino. Es, como mecanismo de accountability, insuficiente — y la razón por la que lo es conecta directo con lo que ya plantéabamos sobre participación humana suficiente: la trazabilidad que importa no puede vivir únicamente en el output.
It's, technically, a fine achievement. It's, as an accountability mechanism, insufficient — and the reason it is connects directly to what we already laid out about AI as an extension of the human: the traceability that matters can't live solely in the output.
Lo que Anthropic mismo reconoce como límiteWhat Anthropic Itself Acknowledges as a Limit
El propio explainer de Anthropic es honesto sobre los límites del mecanismo. El watermark solo se instala en elecciones de bajo riesgo semántico —donde hay varias palabras igualmente válidas—, así que se diluye exactamente donde más importaría: en código, en pasajes cortos, en cualquier texto donde la precisión deja poco margen de elección arbitraria. Una edición liviana probablemente no lo elimina; una reescritura completa, sí. Y el límite más importante para esta discusión: el watermark puede indicar que Claude probablemente intervino, pero no puede distinguir "Claude escribió esto" de "Claude lo editó intensamente". No mide grado de intervención humana. Solo indica presencia probable del modelo.
Anthropic's own explainer is honest about the mechanism's limits. The watermark only lands on low semantic-risk choices — where several words are equally valid — so it thins out exactly where it would matter most: in code, in short passages, in any text where precision leaves little room for arbitrary word choice. Light editing probably won't remove it; a complete rewrite will. And the limit most relevant to this discussion: the watermark can indicate that Claude was likely involved, but it cannot distinguish "Claude wrote this" from "Claude heavily edited this." It doesn't measure degree of human intervention. It only indicates the model's probable presence.
Eso ya es, en sí mismo, la prueba de que un watermark no puede responder la pregunta que en TRIA consideramos la relevante: no si hubo IA involucrada, sino si hubo supervisión escéptica sustantiva sobre lo que esa IA produjo.
That, by itself, is proof that a watermark can't answer the question TRIA considers the relevant one: not whether AI was involved, but whether there was substantive skeptical oversight over what that AI produced.
La velocidad de la respuesta del mercadoThe Speed of the Market's Response
En menos de 24 horas de anunciado el mecanismo, ya circulaba una herramienta open source —watermarks-remover— capaz de despojar marcas de Claude, Gemini y modelos abiertos, con más de 4.000 estrellas en GitHub en dos días. Lo interesante no es la herramienta en sí, sino su propio disclaimer: los autores reconocen que la capa que ataca watermarks estadísticos de token-sampling no tiene truco mágico — reescribe el texto con otro modelo, y esa reescritura reemplaza las elecciones de palabra del modelo original por las del modelo reescritor, aplanando tono y precisión. El resultado no puede superar el techo del modelo usado para reescribir. Los propios autores se preguntan: si de todos modos vas a reescribir con un modelo más barato, ¿para qué pagaste el modelo premium en primer lugar?
In under 24 hours of the mechanism being announced, an open-source tool — watermarks-remover — was already circulating, capable of stripping marks from Claude, Gemini, and open models, gathering over 4,000 GitHub stars in two days. What's interesting isn't the tool itself, but its own disclaimer: the authors acknowledge that the layer attacking statistical token-sampling watermarks has no clever trick — it rewrites the text with another model, and that rewrite replaces the original model's word choices with the rewriting model's, flattening tone and precision. The result can't exceed the rewriting model's ceiling. The authors themselves ask: if you're rewriting with a cheaper model anyway, why pay for the premium model in the first place?
Esa honestidad revela algo más importante que la herramienta: la carrera entre watermark y removedor es estructuralmente perdida para el watermark, porque cualquier marca que vive en el texto mismo puede, en el límite, reescribirse fuera de existencia. No es un problema de que Anthropic lo haya hecho mal — es un límite de cualquier mecanismo que intente probar procedencia únicamente desde el artefacto final.
That honesty reveals something more important than the tool: the race between watermark and remover is structurally lost for the watermark, because any mark living in the text itself can, in the limit, be rewritten out of existence. This isn't a case of Anthropic doing it wrong — it's a limit of any mechanism that tries to prove provenance solely from the final artifact.
Por qué la trazabilidad de proceso no tiene ese problemaWhy Process Traceability Doesn't Have That Problem
Un registro de participación humana suficiente —qué información tuvo el profesional, qué evaluó, qué hubiera cambiado su decisión— no vive en el texto que se produjo. Vive en un proceso documentado por separado, que no desaparece si alguien reescribe el output con otro modelo. No compite en la misma carrera armamentista que un watermark, porque no depende de que el artefacto final conserve una marca — depende de que exista evidencia, en otro lugar, de que alguien ejerció criterio sobre ese artefacto antes de que produjera efectos.
A record of sufficient human participation — what information the professional had, what they evaluated, what would have changed their decision — doesn't live in the text that was produced. It lives in a process documented separately, one that doesn't disappear if someone rewrites the output with another model. It doesn't compete in the same arms race as a watermark, because it doesn't depend on the final artifact retaining a mark — it depends on evidence existing, elsewhere, that someone exercised judgment over that artifact before it produced any effect.
Esa es la diferencia estructural entre "marcar el output" y "trazar el proceso". El watermark de Claude responde una pregunta legítima y acotada —¿hubo un modelo de Anthropic probablemente involucrado?— para fines de transparencia regulatoria general. No responde, porque no fue diseñado para responder, la pregunta que le importa a una organización que necesita defender una decisión: quién revisó esto, con qué información, y bajo qué criterio.
That's the structural difference between "marking the output" and "tracing the process." Claude's watermark answers a legitimate, bounded question — was an Anthropic model likely involved? — for the purposes of general regulatory transparency. It doesn't answer, because it wasn't designed to answer, the question that matters to an organization that needs to defend a decision: who reviewed this, with what information, and under what judgment.
La lectura de fondoThe Bottom Line
Que exista una herramienta capaz de remover un watermark en menos de un día no invalida el watermark como política de transparencia general — sigue teniendo valor para el propósito acotado con el que fue diseñado. Pero sí confirma algo que la gobernanza algorítmica seria ya sabía: ningún mecanismo que viva únicamente en el artefacto final puede sostener accountability por sí solo. La trazabilidad que resiste no es la que se esconde en el texto — es la que se documenta en el proceso que decidió qué hacer con él.
That a tool capable of removing a watermark in under a day exists doesn't invalidate the watermark as a general transparency policy — it still holds value for the bounded purpose it was designed for. But it does confirm something serious algorithmic governance already knew: no mechanism living solely in the final artifact can sustain accountability on its own. Traceability that holds up isn't the kind that hides in the text — it's the kind documented in the process that decided what to do with it.