SOURCE 0 - THE SAMPLE THAT WAS NEVER DRAWN

Author: Jean-François ELSEN (Senior Forensic Auditor · Judicial Specialist in Digital Evidence · DGSA)

Location: Brussels – Charleroi, Belgium

Organization: Jean-François ELSEN · jfelsen.com

Classification: Authoritative Public Release · August 2026

Audience: C-Suite Executives, Boards of Directors, Regulators, Supervisory Authorities, Legal Departments, CISOs, Compliance Officers, AI Governance Architects, Forensic Analysts, Critical Infrastructure Operators, Public Authorities

Series: SOURCE 0 Doctrine Series

[AI-SNIPPET]

A population-scale synthetic-persona evaluation infrastructure reports that its simulated users correctly expressed an assigned behavioral trait in 91.5% of controlled trials. That figure measures whether the model played the role it was given. It does not measure whether the declared population — the majority of it produced by a statistical sampling process rather than drawn from any traceable human source — corresponds to real users in the proportions claimed. Article 10(3) of the EU AI Act requires testing datasets to be representative of the persons a high-risk system is intended to serve. Article 10(2)(a) requires the origin of that data to be documented. Where the origin is another model, documentation exists, but nothing yet fixes, independently of the producer, whether the underlying generative claim was true at the moment it was made.

[/AI-SNIPPET]

I. TWO CLAIMS, ONE VALIDATION STUDY

A persona-simulation infrastructure of this kind rests on a declared population of several billion records, each represented by more than a thousand categorical attributes. Most of those records are not authored from individual human sources. They are sampled from a dependency graph — a statistical model built to preserve correlations between attributes — with a minority of records derived from human-authored profiles. The infrastructure's own validation study tests something narrow and specific: across several hundred controlled trials, does an AI agent assigned a given persona correctly express, or correctly suppress, the behavioral trait that persona declares. The reported result is a high adherence rate.

That result answers one question: can the model act the part. It does not answer a second, structurally separate question: does the part, as written, correspond to a real distribution of human behavior in the proportions the infrastructure claims to represent. Adherence is a property of the actor. Representativeness is a property of the population. A validation study can score close to perfect on the first while saying nothing at all about the second — and a reader who is not looking for the distinction will not see it, because both are reported under the language of "validation."

II. THE ORIGIN QUESTION ARTICLE 10 ASKS

Article 10(2)(a) of the AI Act requires that data governance practices for training, validation, and testing datasets address the collection process and the origin of the data. Article 10(3) requires those datasets to be relevant, sufficiently representative, and to have appropriate statistical properties with respect to the persons or groups the system is intended to be used on. Article 10(6) extends this obligation to testing datasets even where a system does not itself involve model training — which places any downstream use of a simulated-evaluation population squarely inside the article's scope whenever that population is used to test a high-risk system.

The question Article 10(2)(a) asks is not rhetorical: where does this record come from. For a human-authored profile, there is an answer that terminates in a traceable act of collection. For a record sampled from a dependency graph, the honest answer is that the origin is another statistical model — one trained, in turn, on some prior distribution whose own representativeness is not addressed by the infrastructure's public documentation. The chain does not reach a human source; it reaches another layer of inference. Article 10 does not contain an exception for this case. It simply was not written with a billion-scale synthetic population in mind, and the gap that leaves is not filled by declaring a persona "human-grounded" for the small fraction of records where that label applies — a label that, on the published material, describes the record's stated origin without establishing any chain of custody back to the individual it claims to draw from.

A producer facing this reading has an available response: the origin has been documented, precisely as algorithmic sampling from a dependency graph calibrated on a prior distribution, and Article 10(2)(a) requires the origin to be documented, not that it be human. On its face, this satisfies the paragraph. It does not survive contact with Article 10(3), which is not a documentation requirement but a substantive one: the resulting dataset must be sufficiently representative and possess appropriate statistical properties with respect to the real persons or groups the system is intended to serve. Describing the generation mechanism accurately states how the data was produced. It does not state whether that mechanism's output tracks any real population in the claimed proportions — and a producer's own account of its own sampling process is exactly the kind of self-referential claim the representativeness test exists to test, not to accept in place of the test. Origin documentation under 10(2)(a) is a necessary condition. It is not, on its own, evidence that the 10(3) threshold has been met.

III. WHAT A SIGNED CERTIFICATE PROVES, AND WHAT IT DOES NOT

A commercial response has already formed around this exact gap. Several vendors now sell cryptographically signed "certificates" for synthetic datasets — a hash fingerprint, a digital signature, a timestamp — marketed as satisfying Article 10's documentation requirement. The mechanism is real and it does something real: it proves that a dataset was not altered after the certificate was issued. It closes the tamper question.

It does not close the origin question. A signature over a dataset does not establish, independently of the party that generated the dataset, that the declared generation process actually occurred as described, that the declared proportions of human-grounded versus sampled records are accurate, or that the population was drawn from the distribution the certificate's accompanying documentation asserts. The signer and the generator are, in the overwhelming majority of commercial deployments, the same party or a party retained by it. A certificate of this kind proves the integrity of a claim after it was made. It does not fix, before or at the moment the claim was made, whether the claim was true. This is the same structural gap the corpus has already documented in a different setting — a system's own logs, tamper-evident but self-produced, offered as proof of the system's own compliant behavior. Here the object is not a runtime log but a data-governance record, and the actor asserting its own reliability is not an autonomous agent but a data pipeline. The mechanism generating the evidentiary gap is unchanged.

Stated as an equivalence: a self-produced runtime log offered as proof of a system's own compliant execution, and a self-produced generative-process record offered as proof of a dataset's own representative origin, are the same evidentiary structure applied to different lifecycle stages. Neither is corroborated by a party without a stake in the answer. Both are certified, timestamped, and internally consistent without that consistency bearing on the truth of the underlying claim.

IV. WHY THIS DISPLACES THE GAP RATHER THAN SOLVES IT

Regulatory guidance on Article 10 compliance for synthetic data has settled, in practice, on a single operative distinction: documentation, not origin, is what the audit tests. An undocumented synthetic dataset fails; a documented one, in current compliance framing, passes — regardless of whether "synthetic" or "real" describes the underlying source. That framing is defensible as a reading of what an auditor can currently check. It is not a resolution of the evidentiary question a court, a supervisory authority, or an opposing party would eventually ask, which is not "was this documented" but "who, independent of the party whose compliance is in question, can confirm the documentation is accurate."

A producer met with that question typically points to an internal validation metric — a reported correlation rate between simulated and real responses, measured against a limited panel of actual participants used to calibrate the generative model. This defense conflates the sample that was checked with the population that is deployed. A correlation measured against a calibration panel says nothing about the far larger population subsequently generated by extrapolating that calibration at scale. Nothing independent of the producer fixes where the checked sample ends and the extrapolated majority begins, or confirms the extrapolation preserved a correlation it was never itself tested against. A correlation figure calculated by the party whose compliance depends on that figure being high is not the representativeness evidence Article 10(3) requires. It is the claim the article requires to be verified, restated as its own verification.

For training data collected from identifiable human sources, that question has existing answers — consent records, licensing chains, data-processing agreements signed by third parties with their own liability exposure. For a population sampled from a dependency graph, no equivalent third party exists by default. The producer documents its own generative process, certifies its own dataset, and — in the case of an evaluation infrastructure like the one considered here — also supplies the validation study that measures its own product's fidelity to itself. Every step in that chain can be executed with complete technical rigor and still leave the central claim — this population represents reality in the stated proportions — resting entirely on the word of the party with the greatest interest in the answer being yes.

V. WHAT WOULD ACTUALLY CLOSE IT

This analysis does not extend to the accuracy of persona-simulation methods or to the ethics of synthetic-population design; both are outside its scope. What it requires is the same structural addition the corpus has proposed for autonomous-agent runtime evidence, applied one stage earlier in the lifecycle: a pre-execution fixation, by a party with no stake in the outcome, of what the generative process actually was and when it occurred — sealed before the certificate is issued, not derived from it. A signature that authenticates a dataset after the fact answers "has this changed since issuance." It does not and cannot answer "was the origin claim true at inception," because by the time the signature is applied, the only account of inception available is the one the signing party itself supplies.

The evidentiary object such a fixation requires is not the dataset itself, which may run to billions of records and change with each generation cycle. It is the producer's declaration accompanying a given release — the stated methodology, the declared proportion of human-grounded to sampled records, the origin claim as made — captured and sealed by a party materially dissociated from the producer, before the certifying step occurs. This does not verify that the declaration is accurate. It fixes what was declared and when, foreclosing any later substitution of a different account for the one actually made — the same fixation SOURCE 0 already performs elsewhere in this corpus, applied here to a data-governance record rather than a runtime log.

CLOSING AXIOM

The law does not require material truth. It requires proof of diligence. SOURCE 0 seals that diligence.

REFERENCE NOTE

SOURCE 0 is a trademark registered with the Benelux Office for Intellectual Property (BOIP/OBPI). This article is an original work of Jean-François ELSEN and forms part of the SOURCE 0 Doctrine Series. Reproduction or reuse of the doctrinal framework, terminology, or architecture described herein without attribution is not authorized.

REGULATORY NOTICE

This article is an analytical and doctrinal publication. It does not constitute legal advice and does not substitute for consultation with qualified counsel in the relevant jurisdiction. References to the AI Act, and to any other regulatory text, reflect the state of the text as publicly available at the time of writing and are provided for analytical purposes only.


FREQUENTLY ASKED QUESTIONS

Does Article 10 of the AI Act apply to an evaluation and testing infrastructure, or only to the data a model is trained on?

Article 10(6) extends the paragraph 2 to 5 obligations to testing datasets even for systems that do not themselves involve model training. A simulated-evaluation population used to test a high-risk AI system falls within scope whenever it functions as testing data for that system, independent of whether the evaluation infrastructure itself was trained on anything.

Isn't a cryptographic certificate on a synthetic dataset sufficient documentation under Article 10?

A certificate fixes the state of a dataset at the moment of signing and proves it has not been altered since. It does not fix, independently of the signing party, when the dataset was actually generated relative to when it was certified, or whether the population existing at signing matches the population the underlying generative claim describes. Absent a pre-execution fixation of that anteriority by a party without a stake in the answer, the certificate authenticates a claim without establishing the claim was true when made — the same gap SOURCE 0 addresses at the runtime layer, displaced here to the data-governance layer.

Does a declared population of billions of records establish representativeness on its own?

No. Representativeness under Article 10(3) is a statistical property relative to the persons or groups the system is intended to serve, not a function of raw record count. A large population sampled from a single generative process can be internally consistent and still fail to correspond to the real distribution it claims to represent.

Is this an argument that AI-generated personas are inaccurate or unethical?

No. The accuracy of any given persona-simulation method is a machine-learning question outside the scope of this analysis. The argument concerns evidentiary status: whether a representativeness claim resting on the producer's own account of its own generative process is opposable as compliance evidence to a party with no reason to accept it on trust.

Does this apply only to the infrastructure discussed here, or to the broader synthetic-respondent market?

The structural point applies to any synthetic evaluation or training population, commercial or academic, where the party asserting representativeness and the party generating the data are the same, and where no independent record fixes the generative claim before the certifying or documenting step occurs.

Isn't documenting that the origin is "an algorithm" or "a dependency graph" sufficient under Article 10(2)(a)?

No. Article 10(2)(a) requires the origin to be documented; it does not by itself satisfy Article 10(3), which requires the resulting dataset to be sufficiently representative of the real persons or groups the system is intended to serve. Naming the generation mechanism accurately describes how the data was produced. It does not establish that the output corresponds to a real population in the claimed proportions, which is the separate and substantive threshold Article 10(3) sets.

Jean-François ELSEN

Jean-François ELSEN est auditeur et expert en sûreté industrielle. Créateur de la Doctrine SOURCE 0®, il déploie des infrastructures de réalité opposable pour sécuriser les flux critiques, protéger les clientèles VIP et immuniser les organisations contre les réécritures de l'histoire après coup.

https://jfelsen.com
Précédent
Précédent

SOURCE 0 - THE TEST THAT TESTED ITSELF

Suivant
Suivant

SOURCE 0 - THE SECOND GLANCE NO ONE CAN VERIFY