The foundational models that underpin generative AI technologies feed off vast, almost unfathomable quantities of data. Better data usually means better models — and the path to AGI depends on a steady supply of original human-created content. Yet that supply isn’t keeping pace with demand.
Leaders, including Elon Musk and Goldman Sachs chief data officer Neema Raphael, have hinted that human data is running low, while Stanford University’s 2026 AI Index Report estimates that it will be exhausted within the next six years.
To keep pace, AI firms such as Anthropic are turning to artificially-generated data and content (known as synthetic data) to train generative AI models. This has clear advantages: it’s fast, cost-effective, and has been effective at targeting gaps in verifiable tasks, such as mathematics or coding. Used alongside real data, it can enhance the dataset and mitigate its weaknesses.
At the same time, Reinforcement Learning from Human Feedback (RLHF) has long been used to align the outputs of LLMs with human goals, to tune up results and sift out unacceptable and useless language and imagery. RLHF is typically conducted via low-paid human labor.
“The more immediate and measurable risks of contaminated or synthetic-heavy data are, Lee points out, loss of accuracy and the increased danger of amplifying serious bias, such as sexist or homophobic outputs. Long-term, the data ouroburos could create systems with self-referential goals.”
Significant cracks are now appearing in RLHF as whistleblowers report that some workers use chatbots to complete tasks faster and stave off the boredom of their ‘microwork.’ This AI inbreeding produces something akin to a copy of a copy, the equivalent of a blurry JPEG, screenshotted again and again.
Copies of copies
Recent research finds that if a model is offered only recursive data from one source, ‘model collapse’ is inevitable. Rather than an instantaneous failure of the AI, as the term might suggest, it’s a degenerative process, culminating in distorted or nonsensical outputs. Being trained on polluted data, the models then misperceive reality. It all gets rather murky here as the signs are subtle and developers might not even realize it’s happening.
“The difficulty is being able to predict and spot it, because all you know is that the model isn’t working as well as it could be,” says Mark Lee, professor of artificial intelligence at the University of Birmingham.
The more immediate and measurable risks of contaminated or synthetic-heavy data are, Lee points out, loss of accuracy and the increased danger of amplifying serious bias, such as sexist or homophobic outputs. Long-term, the data ouroburos could create systems with self-referential goals.
As Dani Shanley, philosophy professor and senior researcher at the Brightlands Institute for Smart Societies (BISS) explains, synthetic data removes any direct link to real-world events or people, making these underlying decisions even harder to comprehend. “This places unprecedented power in the hands of developers, who can shape reality through data design while potentially obscuring the constructed nature of these representations behind claims of algorithmic objectivity,” she writes for the Ada Lovelace Institute.
One proposed solution is stricter verification of data origins, with a premium on clean, diverse, human-created datasets. Lee likens this to food labeling: developers and users will want transparency about how AI training data is sourced and processed. “But companies may still obscure data provenance, if they’re using the equivalent of battery-farmed AI,” he says.
For Anelia Kurteva, Assistant Professor in Data Management for AI at the University of Birmingham, the future depends on clear records of data origins and transformations that enable accountability, support human agency, and make it easier to understand and govern system behavior. “In the AI scaling race, the pursuit is the fastest result with the highest quality answer,” she says. “But with bigger models, we are losing track of data, decision-making — everything.”
The Alan Turing Institute recommends that teams justify the use of synthetic data before generation, and set out how it will be stored and evaluated. Provenance, shaped around human goals and agency, must come first and be documented clearly enough to audit and review.
Kurteva also argues that responsible AI depends on multidisciplinary collaboration, not just technical optimization. Ethical data management, privacy, consent, and security must be treated as core design issues. “That requires closer collaboration between engineers, legal experts, ethicists and computer scientists to build systems that we can trust,” she says. “A frontier lab composed only of engineers will hurt us all.”
Synthetic data may bring AGI closer into view. But without clarity about where data comes from and how it has been transformed, trust is under threat. We have already seen how the spread of misinformation can destabilize public discourse and undermine confidence in shared sources of truth. As AI systems become larger and more autonomous, data provenance is no longer a technical nice-to-have — it is the basis of trustworthy AI. Without it, confidence in AI’s outputs, and ultimately in the systems themselves, begins to collapse.
Ethical concerns addressed in this article:
- What mechanisms will allow AGIs to practice transparency about their own limitations, origins, authorship, and biases, ensuring their “self-understanding” is clear to human collaborators?
- How does the AGI avoid bias, discrimination, and reinforcement of prejudices in the data, algorithms, and outcomes it produces?
- How do AGI’s processes for impact and risk assessment integrate multidisciplinary, multicultural, and stakeholder perspectives?

Megan Carnegie is a London-based independent journalist who reports on technology, work, and business for publications like Fast Company, Business Insider, and BBC. Her work is underpinned by a desire to investigate what’s not working in the working world and how more equitable conditions can be secured for workers. You can reach her at megancarnegiejournalist@gmail.

