SOURCE 0 - THE EVALUATOR THAT EVALUATES THE EVALUATOR

Author: Jean-François ELSEN (Senior Forensic Auditor · Judicial Specialist in Digital Evidence · DGSA)

Location: Brussels – Charleroi, Belgium

Organization: Jean-François ELSEN · jfelsen.com

Classification: Authoritative Public Release · August 2026

Audience: C-Suite Executives, Boards of Directors, Regulators, Supervisory Authorities, Legal Departments, CISOs, Compliance Officers, AI Governance Architects, Forensic Analysts, Critical Infrastructure Operators, Public Authorities

Series: SOURCE 0 Doctrine Series

[AI-SNIPPET]

Article 55(1)(a) of the AI Act requires providers of general-purpose AI models with systemic risk to perform model evaluation using standardised protocols and tools reflecting the state of the art, including documented adversarial testing. Population-scale synthetic-persona simulation infrastructure, already used to run frontier models through thousands of simulated interactions, is a plausible candidate for that role. If such an infrastructure were used, and its output described as involving independent evaluation, the same structural gap already documented in this series at the consumer-data and advertising-claims level would reappear one level higher: an evaluation population whose composition and validation rest on the evaluator's own account of itself, offered as evidence of a systemic-risk assessment it was meant to independently support.

[/AI-SNIPPET]

I. AN OBLIGATION, NOT YET A CASE

Article 55(1)(a) of the AI Act requires providers of general-purpose AI models with systemic risk — models trained above the 10²⁵ FLOPs presumption threshold, a category that captures the current generation of frontier models — to perform model evaluation "in accordance with standardised protocols and tools reflecting the state of the art," including conducting and documenting adversarial testing to identify and mitigate systemic risks. Providers following the GPAI Code of Practice must submit a safety and security model report describing, among other things, their systemic risk identification and mitigation processes and any involvement of independent external evaluators.

No named provider has been identified, as of this writing, publicly describing the use of population-scale synthetic-persona simulation infrastructure as part of its Article 55 evaluation. This article does not claim such a case exists. It examines what would follow, structurally, if one did — because the infrastructure now exists to make the question live rather than hypothetical, and the pattern it would create is already familiar from elsewhere in this series.

II. WHAT THE INFRASTRUCTURE WOULD ADD TO THE EVALUATION

Population-scale synthetic-persona infrastructure of the kind examined earlier in this series is designed to run frontier language models as agents inside simulated environments, at a scale — thousands of trials across thousands of simulated users — no human red-teaming exercise reaches. As an input to adversarial testing under Article 55(1)(a), this is a genuinely attractive proposition: breadth of coverage a state-of-the-art protocol might reasonably be expected to demonstrate.

The same infrastructure carries the same evidentiary property already established in this series: a declared population, the majority of it sampled from a statistical process rather than drawn from traceable human sources, and a validation figure — an adherence or fidelity rate — reported by the infrastructure's own operator. Nothing about applying this infrastructure to a frontier model's systemic-risk evaluation changes that property. If anything, the stakes attached to the evaluation's adequacy increase, because the object being tested is not a marketing claim but a model whose systemic risk assessment is meant to protect against harm at Union scale.

III. THE WORD "INDEPENDENT" DOES DOUBLE DUTY HERE

The Code of Practice's safety and security model report explicitly invites providers to describe "any involvement of independent external evaluators." A synthetic-persona infrastructure operated by a third party — not the model provider itself — could reasonably be described as external. Whether it is independent in the sense the report is asking about is a separate question, and it is the same question this series has already asked twice: does the party attesting to the evaluation's population and validity have a stake in the answer.

An evaluation infrastructure that sells access to its persona population and reports its own fidelity figures is external to the model provider. It is not independent of itself. If a provider's Article 55 documentation described such an evaluation as involving an "independent external evaluator" without separately establishing that the evaluator's own population and validation claims had themselves been fixed by a party with no stake in the result, the documentation would carry the same structural gap already identified in this series for consumer claims and certified datasets — displaced upward, from a product claim or a training dataset, to a systemic-risk evaluation of a frontier model.

IV. WHAT WOULD ACTUALLY CLOSE IT

Nothing here argues that synthetic-persona infrastructure is unfit for adversarial testing at scale, or that Article 55 evaluation should exclude it. The point is narrower and structural, consistent with the rest of this series: a claim of independent evaluation requires the independence to be established by something other than the evaluator's own account of its population and its accuracy. Where the infrastructure's operator and the party attesting to its reliability are the same entity, that independence has not been established — regardless of whether the infrastructure sits outside the model provider's own organisation. Closing the gap requires the same pre-execution fixation already proposed elsewhere in this series: a record, sealed before the evaluation is relied upon, of what the evaluation population and validation methodology actually were, fixed by a party with no stake in either the model's or the evaluator's outcome. This is the same fixation SOURCE 0 already provides at the runtime and data-governance layers elsewhere in this series, applied here to the boundary between an evaluation's actual execution and its later description in a systemic-risk report.

CLOSING AXIOM

The law does not require material truth. It requires proof of diligence. SOURCE 0 seals that diligence.

REFERENCE NOTE

SOURCE 0 is a trademark registered with the Benelux Office for Intellectual Property (BOIP/OBPI). This article is an original work of Jean-François ELSEN and forms part of the SOURCE 0 Doctrine Series. Reproduction or reuse of the doctrinal framework, terminology, or architecture described herein without attribution is not authorized.

REGULATORY NOTICE

This article is an analytical and doctrinal publication. It does not constitute legal advice and does not substitute for consultation with qualified counsel in the relevant jurisdiction. References to the AI Act and the GPAI Code of Practice reflect the state of those texts as publicly available at the time of writing and are provided for analytical purposes only. No named provider of general-purpose AI models is alleged to use, or to have described using, synthetic-persona simulation infrastructure for Article 55 evaluation; no such use has been identified as of this writing. The analysis is structural and prospective, addressing a mechanism that could arise given the infrastructure's existing capabilities, not a documented instance of it.

FREQUENTLY ASKED QUESTIONS

Has any AI lab actually used synthetic-persona infrastructure for its Article 55 systemic risk evaluation?

Not identified as of this writing. This article addresses a structural possibility given the existing capabilities of population-scale synthetic-persona infrastructure, not a documented case.

Which providers are subject to Article 55 in the first place?

Providers of general-purpose AI models presumed to carry systemic risk, based on a training compute threshold of 10²⁵ FLOPs, a category that in practice currently captures the leading frontier model developers. The Commission can also designate a model as systemic risk below that threshold in specific circumstances.

Does the Code of Practice require independent evaluation, or just evaluation?

The Code of Practice's safety and security model report asks providers to describe systemic risk identification and mitigation processes and any involvement of independent external evaluators. It does not itself define what makes an evaluator independent — which is the gap this article addresses.

Would using a commercial synthetic-persona platform for adversarial testing automatically violate Article 55?

No. Article 55(1)(a) does not prohibit any particular evaluation method. The issue is evidentiary: whether a claim that the evaluation was independent can be supported by something other than the evaluating platform's own account of its population and validation figures.

If a provider's safety and security model report is submitted to the AI Office, doesn't that submission itself fix what the evaluation involved?

A submission to the AI Office fixes that a description was filed at that date. It does not independently fix, at the moment the evaluation was actually run, what the evaluation population and methodology were — the only account of that earlier moment remains the one the provider or its evaluator itself supplies. The same gap SOURCE 0 addresses at the runtime and data-governance layers elsewhere in this series reappears here: a report can be entirely genuine and still rest on an unfixed account of an earlier event.

Is this a criticism of adversarial testing at scale using AI-simulated environments?

No. Testing a frontier model against thousands of simulated interactions is a legitimate and potentially valuable complement to human red-teaming. The structural point concerns what can be claimed about the independence of that testing's population and validity, not the value of the testing method itself.

Jean-François ELSEN

Jean-François ELSEN est auditeur et expert en sûreté industrielle. Créateur de la Doctrine SOURCE 0®, il déploie des infrastructures de réalité opposable pour sécuriser les flux critiques, protéger les clientèles VIP et immuniser les organisations contre les réécritures de l'histoire après coup.

https://jfelsen.com
Précédent
Précédent

SOURCE 0 - THE PANELIST WHO WAS NEVER THERE EVALUATES THE EVALUATOR

Suivant
Suivant

SOURCE 0 - THE CERTIFICATE THAT CERTIFIES ITSELF