Synthetic Data: What It Is, How It Is Generated, and What It Does Not Guarantee
01

One-sentence definition

Datos generados artificialmente para entrenar, probar o simular escenarios cuando se controla su procedencia y calidad.

02

Definition: data created artificially for a specific purpose

Synthetic data is data generated artificially to reproduce, with varying degrees of fidelity, characteristics or properties of real data that are relevant to a particular use case. The word “synthetic” describes how the data originated; it does not certify its quality, its usefulness for every task, or its anonymity.

The fundamental distinction is between the origin of a dataset and its properties. A synthetic record is not collected directly as an observation of a real-world phenomenon: it is produced through some generation process. For a particular purpose, it may be important for the data to retain certain patterns from a reference dataset, such as relationships between variables or the presence of relevant cases. What needs to be preserved depends on the task; there is no single property that makes every synthetic dataset useful.

The Spanish Data Protection Agency (AEPD) describes synthetic data as data generated artificially and notes that, for a specific use case, it should preserve characteristics and properties of real data. This condition of context-dependent usefulness is not a claim that the result faithfully reproduces every aspect of the original dataset, nor that it carries no privacy risk.

So a statement such as “this dataset is synthetic” answers, at most, a question about its origin. To assess what it can be used for, ask what process generated it, which aspects of the reference data were intended to be preserved, and what checks were performed for the proposed use.

03

How generation and validation are planned

A synthetic-data project starts with its purpose, not with the choice of a tool. A dataset suitable for testing whether software works may be unsuitable for estimating a trend, studying rare cases, or training a perception system. Before generating anything, clarify which decision, test, or research activity the data is intended to support.

The next step is to define which properties matter for that purpose. These might include the structure of the data, relationships between fields, or the presence of examples of particular situations. The criterion should not be “make it look similar” in the abstract: specify what kind of similarity would be useful and how it will be checked. The sources consulted do not identify a universal list of properties or a metric that works for every use.

The generation procedure is then selected and documented. The sources provided confirm that data can be generated artificially and refer to algorithmic generation, but they are not sufficient to establish an exhaustive technical taxonomy or to compare particular methods. When reviewing a dataset, it is therefore more rigorous to ask for a description of the method actually used than to infer it from the label “synthetic” alone.

Finally, assess the result against its intended purpose and examine potential privacy implications separately. A check of statistical similarity, if one has been conducted, does not by itself demonstrate that the dataset is suitable for a specific task; nor does it show that the dataset does not reproduce information from its source. The evaluation should state what was tested, against which reference, and what limitations remain.

Evaluation sequence

This outline organizes the questions to ask; it does not prescribe a single technical method.

  1. 01Specify the intended use and who will use the dataset.
  2. 02Identify the properties the task needs to preserve and those that are not relevant.
  3. 03Document the source of the reference data and the generation procedure.
  4. 04Check usefulness against task-specific criteria, not just a generic measure of similarity.
  5. 05Separately assess whether information from the source dataset could be reproduced or disclosed.
  6. 06Report the checks performed, their results, and known limitations.
04

Three applied examples

The following scenarios illustrate questions a team might ask; they are not claims that a particular dataset has been generated, validated, or deployed in these ways. In every case, usefulness depends on the task and on the available evidence.

blocksAreNotParagraphsPlaceholder

05

What to check: usefulness, coverage, and privacy

Usefulness is not a universal property of a dataset; it is usefulness for a particular task. Ask what activity was validated and whether it matches the proposed use. A dataset that lets a team test whether an application accepts certain formats may not support a conclusion about a model’s accuracy. Likewise, an evaluation result should not be generalized to populations, environments, or decisions that were not examined.

Coverage matters too. The generation process may fail to preserve every relevant case, and a dataset may represent different groups, situations, or combinations of variables with varying degrees of fidelity. A large volume of data, or agreement on a few summary statistics, is not enough. The team should explain which dimensions it compared and which cases were outside the evaluation.

Privacy requires a separate examination. The fact that data were not collected directly as real-world records does not prove that the generation process cannot reproduce or make it possible to infer information from the material used to create them. The sources provided do not report privacy-test results for a specific dataset, or offer a general guarantee that applies to every method. Any privacy claim should therefore be tied to identifiable tests, assumptions, and a context of use.

It is also useful to document the reference dataset, the purpose of generation, the procedure, the evaluations, and the limitations. Without this information, it is not possible to distinguish an evidence-backed claim from a general description. Privacy evaluation also does not replace analysis of the obligations and risks relevant to a particular context: this article is not legal advice.

Questions to ask about different claims

The label alone does not answer these questions.

ClaimWhat to ask forWhat the claim alone does not establish
“It is useful”The task evaluated, the criteria, the reference, and the results.That it will work for other tasks or populations.
“It is representative”Which groups or properties were compared and how they were measured.That every relevant case is covered.
“It protects privacy”Which risk was assessed, using what tests, and under which assumptions.That “synthetic” is equivalent to anonymous.
“It resembles real data”Which dimensions of similarity were measured and why they matter.That it preserves every useful property or reveals no information.
06

Related concepts that should not be confused

Synthetic data and anonymized data describe different things. “Synthetic” refers to artificial origin; “anonymized” is a claim about data processing and the risk of identifying or re-identifying people in a particular context. A synthetic dataset may have been generated from real data, so its artificial nature is not enough to determine what information it retains or what risks it presents. Nor should anonymization be assumed without evidence.

Data augmentation usually means creating variations of existing examples through transformations, with the aim of expanding or modifying the material available for a task. It is not automatically equivalent to generating a dataset through another process. To classify a particular case, it matters whether original examples were transformed, observations were created using rules or simulation, or a model was used; a commercial or informal label does not settle the question.

Simulation generates observations according to rules or a representation of a process. It can produce synthetic data, but the term describes the mechanism or environment used for generation, while “synthetic” describes the artificial origin of the resulting data. How good a simulation is depends on its assumptions and on the relationship between what is simulated and the real task.

Model-generated data are another possible family of procedures, but it should not be assumed that all synthetic data come from a model trained on real records. The institutional sources available here do not document every family, its requirements, or a comparison of performance in sufficient detail. For a specific claim, ask for a description of the method and avoid turning a technical possibility into a universal definition.

Finally, data contamination, generalization, and AI evaluation are related but distinct concepts. Contamination concerns unsuitable information, or information related to an evaluation, affecting the data or the process being evaluated; generalization concerns how a system behaves beyond the examples used to develop it; and evaluation is the process of measuring that behavior. Using synthetic data does not, by itself, resolve any of these issues.

Quick distinctions

Classify a concept by the question it answers, not by an assumed guarantee shared with other concepts.

ConceptMain questionDifference from synthetic data
Synthetic dataHow did the data originate?It was generated artificially; that alone does not determine usefulness or privacy.
Anonymized dataWhat identification risk has been assessed?This is a privacy question, not a sufficient description of origin.
Data augmentationWere examples transformed to expand or vary the material?It describes an operation on examples; it is not necessarily independent generation.
SimulationWhich rules or representation produce the observations?It can be a mechanism for producing synthetic data, but it does not define every method.
07

Limitations and common mistakes

A common mistake is to treat statistical similarity as a complete guarantee. A dataset may approximate some properties of the reference material and differ on others. Without knowing which characteristics were compared, it is not possible to claim that it preserves “the data” in general. The relevant degree of fidelity depends on the question being asked.

Another mistake is to extend the result of a test to a broader use. Data that work for testing a software feature are not necessarily suitable for inference, research, or decisions about people. Before reusing them, review whether the purpose has changed and whether the validity criteria have been justified again.

It is also risky to interpret “contains no real data” as “cannot reveal information.” That conclusion requires examining how the dataset was generated and what tests support the claim. The notes accompanying the available sources do not allow us to declare that any particular method is immune to reproducing information, or to quantify a general level of risk.

Finally, do not assume that a synthetic dataset always replaces real data or eliminates every limitation related to access, bias, or quality. It may help with particular workflows, but relevance and risk must be assessed case by case. If documentation, metrics, or tests are missing, the reasonable conclusion is that the evidence is insufficient—not that the dataset is necessarily useless or necessarily safe.

08

Practical criteria for assessing a claim

When someone provides a dataset or makes a claim about synthetic data, start by asking for a verifiable description: who generated it, what material or reference it was based on, which procedure was used, and what task it was designed for. If those answers are unavailable, it is not possible to assess precisely either the scope of its usefulness or the meaning of a privacy claim.

Next, check whether the evaluation matches the purpose. Ask which properties were measured, which relevant cases were considered, what comparison was used, and which limitations were acknowledged. An isolated metric has no universal meaning: you need to know what it is intended to measure and why it is appropriate for that application.

For privacy, ask for an explanation of the tests performed and the type of risk examined, without accepting “synthetic” as a synonym for “anonymous.” For performance, ask which use was validated and which uses remain outside the evidence. Keep the two evaluations separate: a favorable conclusion in one does not automatically resolve the other.

In short: define the purpose, examine the method, require task-specific validation, and treat privacy as a separate issue. If the references, tests, or assumptions are unknown, record that uncertainty. This caution helps distinguish a documented property from an expectation and prevents the term from being given a broader meaning than it has.

A short checklist

Use these questions to review documentation or speak with the person presenting the dataset.

  1. 01What is the specific use, and which decision or test is the dataset intended to support?
  2. 02Which properties were expected to be preserved, and which were evaluated?
  3. 03How was the dataset generated, and how is it related to the reference data?
  4. 04What evidence supports its usefulness for that task, and what limitations were reported?
  5. 05What separate tests were performed to evaluate privacy?
  6. 06Is the conclusion limited to what was actually tested?
09

Quick examples

10

Related concepts

11

Sources consulted