Population-wide data without the privacy risks? IR Stone/Shutterstock

For the past decade, most public debates about data have revolved around a familiar anxiety: too much of our personal information is being collected from our records and our online activity, and then shared and monetised. That concern has not gone away. But a shift is now under way – towards synthetic data.

Increasingly, data used by governments, hospitals, banks and technology companies to make decisions about the people they serve is not gathered directly from real people and events. It can be synthetic: generated by computer models that learn the statistical patterns in a real dataset and then produce artificial records that follow the same patterns. The aim is to keep the data useful while reducing the risk that individuals can be identified.

In a society concerned about surveillance and misuse of personal information, synthetic data promises useful knowledge in a way that may allay public concerns. In England, for example, researchers can use the Simulacrum, a synthetic cancer dataset built from NHS registry data, without ever seeing real patients’ records. Its artificial patients are generated from groups of at least 50 real cases – but this also means that rarer cancers are represented less realistically.

Building such a dataset does, of course, require access to the real records in the first place: the model has to learn from them before it can imitate them. Synthetic data therefore shifts the privacy question rather than dissolving it – the original data still has to be held, and protected, somewhere.




Read more:
Can baby monitors protect children without
data security risks?



Traditionally, the law has asked whether data has been collected lawfully, used fairly and disclosed properly. When important decisions rely on synthetic data, those questions still matter – but they are no longer sufficient. We also need to know who designed this artificial world – and which populations, risks and behaviour were treated as “normal” when it was built. These questions are central to my research, including my forthcoming book, Corporate and Financial Information in the Age of Synthetic Data.

In other words, synthetic data does not take human judgment out of the picture. It moves it to an earlier, less visible stage. The crucial choices – what counts as typical, what can safely be left out – are made when the dataset is built, long before anyone uses it to make a decision about us.

An illusion of objective data?

That shift extends across healthcare, welfare, policing and finance, where synthetic data may shape how institutions define disease, vulnerability, risk and suspicious behaviour. Synthetic transaction data, for instance, is being used to develop and test the systems that flag possible fraud and money laundering – systems that can help decide whether a payment goes through or an account is frozen.

A report published by the Financial Conduct Authority, the UK’s financial regulator, based on the work of an expert group it set up, stresses that firms need to assess and mitigate the risks, including bias.

The wider risk to society is subtler than a privacy breach, but serious nonetheless. Synthetic data can create an illusion of objectivity because it looks scientific and orderly. Yet it is still the product of human choices, and those choices can carry bias.

If the generating model is too blunt, it may erase minority experiences or exceptional cases. If several institutions rely on similar synthetic environments, they may converge on the same blind spots.

This is why synthetic data should not be understood simply as a privacy technology, but as part of the emerging politics of knowledge. States and markets depend on systems that reveal where risks lie, where needs are greatest, where money is flowing and where intervention is justified.

As data becomes partly synthetic, the legitimacy of the growing number of decisions made by governments and firms will depend on whether the public can trust the methods they use.

It comes down to governance. Institutions should be able to explain in an accessible way why synthetic data was used, what it was likely to distort and how it was validated. They should also explain what safeguards exist if it proves misleading. That is especially important when the data informs decisions that affect rights, opportunities or how public money is spent.

England’s 2020 exam results show what can go wrong. With exams cancelled during the COVID pandemic, an algorithm brought teacher-assessed A-level grades into line with schools’ past results. Nearly two in five grades were marked down, and high-performing students in disadvantaged areas were hit hardest. The government backed down days later and scrapped the algorithm’s grades.

That system did not use synthetic data, but the lesson carries over. When a model’s version of the world shapes people’s futures, those affected need to know how it was built – and have a real way to challenge it.

schoolboy seen from behind climbing stairs to a footbridge carrying a rucksack and a sports bag.
Data helps steer the decisions that shape futures – but synthetic data comes with challenges.
Mounir Taha/Shutterstock

This also suggests that the language of transparency needs to evolve. For years, data governance has often focused on who has the data, who may share it and who may inspect it. Synthetic data complicates that picture. So attention needs to shift to whoever builds the model that generates the data.

The Information Commissioner’s Office treats synthetic data as part of a broader family of privacy-enhancing technologies. But where synthetic data begins shaping public and economic decisions, governance must also ask who is responsible for how institutions come to know what they claim to know.

These questions matter to every citizen because synthetic data brings both promise and danger. Properly governed, it could help tackle privacy issues and widen access to information that helps society. Poorly governed, it could create systems that are cleaner, safer and more efficient on paper – but detached from the messy truths of real life.

The governance of synthetic data, then, is not a niche issue for engineers and lawyers. It belongs squarely within the larger political conversation about trust, accountability and democratic control in data-driven societies. The question used to be who owned the data. The new question is more unsettling: who gets to manufacture credible versions of the world?

The real challenge is to ensure that as artificial information becomes woven into the fabric of governance, medicine, finance and public administration, we do not unquestioningly allow it to pass for the truth. Synthetic data can be genuinely useful, but we must be able to trust the world it constructs.

The Conversation

Maria Lucia Passador does not work for, consult, own shares in or receive funding from any company or organisation that would benefit from this article, and has disclosed no relevant affiliations beyond their academic appointment.

Leave a Reply

Your email address will not be published. Required fields are marked *