Synthetic Data

Synthetic data are artificially generated datasets that reproduce the structure and statistical properties of real data and are meant to work without the actual original records. Whether a given synthetic dataset really carries no personal reference any more has to be assessed separately in each case. AI models generate text, image or tabular data of this kind to train or test other AI systems, above all where real data is scarce, expensive or hard to use for data protection reasons.

In practice

Synthetic data helps where real training data is missing or legally sensitive, for instance with medical or financial records, and it is being used more and more as high-quality real training data on the open web becomes scarcer. One known risk is model collapse: if a model is trained repeatedly and largely on AI-generated rather than real data, errors and biases can amplify each other instead of evening out. Serious providers mix synthetic data with curated real data and keep checking results against genuine reference data.

Sources

← Back to the glossary