← Publicaciones← Publications
AnálisisAnálisisAuditoríaAuditoríaGobernanza IAGobernanza IATRIATRIA

Que un algoritmo "parezca maduro" no alcanza: por qué la gobernanza de IA necesita evidencia, no declaracionesAn algorithm "looking mature" isn't enough: why AI governance needs evidence, not declarations

Analista de Transfer Pricing. Investigador en gobernanza algorítmica aplicada al cumplimiento fiscal internacional. Creador del Sistema TRIA.Transfer Pricing Analyst. Researcher in algorithmic governance applied to international tax compliance. Creator of Sistema TRIA.

Ulises Lento6 de agosto de 20266 de agosto de 2026~5 min de lectura~5 min read

Cuando una organización audita sus estados contables, nadie se conforma con que el gerente de finanzas diga "está todo bien". Se pide evidencia: papeles de trabajo, muestras representativas, controles cruzados. La auditoría financiera dejó de ser una cuestión de confianza hace más de un siglo.When an organization audits its financial statements, no one settles for the finance manager saying "everything's fine." Evidence is required: working papers, representative samples, cross-checks. Financial auditing stopped being a matter of trust over a century ago.

La pregunta que nadie se haceThe question no one asks

La gobernanza algorítmica todavía no llegó a ese punto.

Algorithmic governance hasn't gotten there yet.

Hoy, cuando una empresa o un organismo público asegura que su sistema de IA es "maduro" —que está bien gobernado, que fue probado, que es confiable— en la mayoría de los casos esa afirmación se sostiene sobre un checklist. Se marcan casilleros: ¿tiene política de uso de IA? Sí. ¿Tiene un responsable designado? Sí. ¿Se testeó antes de implementarse? Sí. Y con eso, se declara la madurez.

Today, when a company or a public agency claims its AI system is "mature" — that it's well-governed, that it was tested, that it's reliable — in most cases that claim rests on a checklist. Boxes get checked: does it have an AI use policy? Yes. Does it have a designated owner? Yes. Was it tested before deployment? Yes. And with that, maturity is declared.

El problema es el mismo que tenía la auditoría antes de exigir evidencia: un checklist mide si algo existe, no si funciona.

The problem is the same one auditing had before it started requiring evidence: a checklist measures whether something exists, not whether it works.

Supongamos que un sistema de scoring —crediticio, aduanero, tributario— toma decisiones sobre miles de casos por día. El checklist dice que el sistema fue auditado. Pero nadie preguntó algo más simple: si dos auditores distintos revisaran el mismo sistema, ¿llegarían a la misma conclusión sobre su nivel de madurez?

Suppose a scoring system — credit, customs, tax — makes decisions on thousands of cases a day. The checklist says the system was audited. But no one asked something simpler: if two different auditors reviewed the same system, would they reach the same conclusion about its maturity level?

Si la respuesta es "no necesariamente", el problema no es el sistema de IA. Es el instrumento con el que lo estamos midiendo.

If the answer is "not necessarily," the problem isn't the AI system. It's the instrument we're measuring it with.

Esto tiene un nombre en cualquier disciplina que mide cosas de forma sistemática: fiabilidad. Una balanza que da un peso distinto cada vez que pesás el mismo objeto no es una balanza confiable, por más precisa que parezca a simple vista. Un instrumento de evaluación de madurez algorítmica que depende de quién lo aplique tiene el mismo defecto.

This has a name in any discipline that measures things systematically: reliability. A scale that gives a different weight every time you weigh the same object isn't a reliable scale, no matter how precise it looks at first glance. An algorithmic maturity assessment instrument whose results depend on who applies it has the same flaw.

Tratar distinto lo que es igualTreating equal things differently

Hay una segunda pregunta que el checklist tampoco responde: ¿el sistema trata de forma distinta a dos casos que deberían tratarse igual?

There's a second question the checklist doesn't answer either: does the system treat two cases differently when they should be treated the same?

En auditoría financiera esto se llama consistencia de criterio. En gobernanza algorítmica se llama sesgo, pero la lógica de fondo es idéntica: si dos solicitudes de crédito son iguales en todo excepto en una variable que no debería pesar en la decisión —el origen del solicitante, por ejemplo— y el sistema las resuelve distinto, hay un problema de gobierno, no un detalle técnico menor.

In financial auditing this is called consistency of criteria. In algorithmic governance it's called bias, but the underlying logic is identical: if two credit applications are identical in every respect except one variable that shouldn't factor into the decision — the applicant's origin, for instance — and the system resolves them differently, that's a governance problem, not a minor technical detail.

Detectar esto no requiere que el equipo de auditoría entienda machine learning. Requiere comparar resultados entre grupos que deberían comportarse igual y ver si efectivamente lo hacen. Es la misma lógica que se usa para detectar anomalías en cualquier proceso: identificar dónde algo se aparta de lo esperado y preguntar por qué.

Catching this doesn't require the audit team to understand machine learning. It requires comparing outcomes across groups that should behave the same way and checking whether they actually do. It's the same logic used to detect anomalies in any process: identify where something deviates from what's expected, and ask why.

No hace falta revisar todoYou don't need to review everything

La tercera idea, y quizás la más práctica: nadie audita el 100% de las transacciones de una empresa. Se define una muestra representativa, se la revisa en profundidad, y las conclusiones se extienden con un margen de confianza conocido.

The third idea, and perhaps the most practical one: no one audits 100% of a company's transactions. A representative sample is defined, reviewed in depth, and the conclusions are extended with a known confidence margin.

Con la IA pasa algo parecido y muy pocos lo están aplicando. No es necesario —ni posible, en sistemas que procesan millones de decisiones— revisar cada output de un algoritmo para saber si está funcionando como se espera. Se puede diseñar una muestra, revisarla con rigor, y saber con qué nivel de confianza esa muestra representa al total. La diferencia entre "revisamos algunos casos al azar" y "diseñamos una muestra estadísticamente representativa" es la diferencia entre una opinión y una auditoría.

Something similar applies to AI, and very few are doing it. It isn't necessary — or even possible, in systems processing millions of decisions — to review every output of an algorithm to know whether it's working as expected. You can design a sample, review it rigorously, and know with what confidence level that sample represents the whole. The difference between "we checked a few random cases" and "we designed a statistically representative sample" is the difference between an opinion and an audit.

De la declaración a la evidenciaFrom declaration to evidence

Estas tres ideas —fiabilidad del instrumento, consistencia entre casos comparables, y muestreo representativo— no son técnicas de inteligencia artificial. Son herramientas de auditoría que existen hace décadas y que la gobernanza algorítmica todavía no terminó de incorporar.

These three ideas — instrument reliability, consistency across comparable cases, and representative sampling — aren't artificial intelligence techniques. They're auditing tools that have existed for decades, and that algorithmic governance still hasn't fully adopted.

Es exactamente el espíritu detrás de la Matriz Diagnóstica de Madurez Algorítmica (MDMA): no alcanza con que una organización declare en qué nivel de madurez está. Hace falta que ese nivel se pueda sostener con evidencia, de la misma manera que un balance no se sostiene con la palabra del contador sino con los papeles de trabajo que hay detrás.

That's exactly the spirit behind the Algorithmic Maturity Diagnostic Matrix (MDMA): it isn't enough for an organization to declare its maturity level. That level has to be sustainable with evidence, the same way a balance sheet isn't sustained by the accountant's word but by the working papers behind it.

La trazabilidad que propone TRIA parte de esa misma premisa: cada decisión algorítmica debería poder reconstruirse y explicarse, no solo declararse conforme. Cuando eso se combina con un instrumento de medición fiable, consistente y verificable por muestreo, la madurez algorítmica deja de ser una autopercepción de la organización y pasa a ser algo que un tercero puede verificar. Que es, en definitiva, lo único que hace que una auditoría valga como tal.

The traceability TRIA proposes starts from that same premise: every algorithmic decision should be reconstructible and explainable, not merely declared compliant. When that's combined with a measurement instrument that's reliable, consistent, and verifiable by sampling, algorithmic maturity stops being an organization's self-perception and becomes something a third party can verify. Which is, ultimately, the only thing that makes an audit worth the name.

ReferenciasReferences

Lento U. (2026). Marco TRIA™: Trazabilidad, Razonamiento, IA, Auditoría. Dirección Nacional del Derecho de Autor, Legajo N.° RL-2026-46255181.

Lento U. (2026). Marco TRIA™: Trazabilidad, Razonamiento, IA, Auditoría. Dirección Nacional del Derecho de Autor, Legajo N.° RL-2026-46255181.

Lento U. (2026). Matriz Diagnóstica de Madurez Algorítmica (MDMA). Dirección Nacional del Derecho de Autor, N.° PV_2026-67047663-APN-DNDA#MJ. Expediente EX-2026-677047649-APN-DNDA#MJ.

Lento U. (2026). Matriz Diagnóstica de Madurez Algorítmica (MDMA). Dirección Nacional del Derecho de Autor, N.° PV_2026-67047663-APN-DNDA#MJ. Expediente EX-2026-677047649-APN-DNDA#MJ.

La automatización sin trazabilidad no es eficiencia. Es riesgo diferido.
Automation without traceability is not efficiency. It is deferred risk.
Compartir