The ethical importance of accurate data: it’s only FAIR

By Amy Lyall

Photo by Bob Aglow

My background as a writer lies in the biological sciences. In this field, the advent of AI has had an impact which simply cannot be overstated. In the decade I have been working in (and writing about) biology I have watched the entire field change. 

AI has applications in some of the most intractable questions life science researchers face. It has changed the face of genetic sequencing, drug discovery, ecological monitoring, and large-scale sample analysis. 

An example: proteins, the active products of gene transcription. These are the molecules which physically carry out the instructions blueprinted in our DNA. Some endure a lifetime. Some are gone in the blink of an eye. And understanding their structure is vital to understanding how they do what they do. 

Over the last five decades, entire careers have been devoted to solving protein structures. But predicting a protein’s three-dimensional structure from its DNA sequence and amino acids can take a human researcher months, even years. Out of billions of proteins, we had structures for approximately 170,000 – each the result of long hours of painstaking work. 

Google DeepMind’s AlphaFold solves proteins in minutes, and does so to near-experimental accuracy. We now have a million structures at our fingertips, and the number is growing every day. In 2024, AlphaFold won the Nobel Prize for Chemistry.

Tools like this have been coupled with massive advances in sequencing. It took thirteen years, thousands of scientists worldwide, and three billion dollars to complete the first human genome sequence in 2003. Today, a full genome can be sequenced in hours for under $600. 

For many years, the main issue facing biologists was the struggle to get the data they needed. Now, the problem is different. Researchers have access to vast quantities of data. The difficulty is in searching it to isolate what they need.

This is where tools like AI can help. Vast amounts of genomic, proteomic, metabolomic, and imaging data are generated every day, far more than can be analysed by traditional methods. Properly trained AI systems can analyse this mountain of information with speed and accuracy, picking up patterns and details too subtle for humans to see.  

And an AGI could be even more transformative. It would give us systems capable of general reasoning, creative scientific insight, and global-scale data analysis. 

In biology, AGI could autonomously generate hypotheses and design its own experiments. It could integrate massive amounts of complex data, from sources all over the world, and achieve a systems-level understanding of life. It could discover novel therapeutics, biosynthetic pathways, and sustainable biotechnologies. It could, essentially, change both our perceptions of the natural world and the tools we use to study it.

However, AGI, like AI (and like humans) can only make accurate decisions with accurate data. Without accessible, well-annotated, and ethically governed information, AGI could make inaccurate or unsafe inferences. 

Biological systems are inherently unpredictable. Working on viruses or genetic manipulation without the correct contextual information could have catastrophic consequences. 

So how do we ensure that the data foundations we are building AI and AGI on are sound?

Scientists worldwide use different methods to generate, label, store, and share data and metadata. Standards are inconsistent, incentives for responsible data sharing are limited, and institutional data silos persist. There is currently no universal approach to storing, searching and sharing scholarly data. 

This means results can’t always be accessed or interpreted, they don’t always have clear and traceable origins, and as a result they are not easy to use or authenticate. 

This data is not suitable as fodder for AI and AGI. If it can’t be traced or authenticated, we don’t know whether it’s inaccurate or incomplete. We run the risk of feeding misinformation into the system. Addressing these issues means urgent implementation of transparent, standardised, and ethically managed data ecosystems. 

A proposal for FAIR research data was published in Nature in 2016 by a consortium of life science researchers, industry professionals, publishers and academics. It sets out four guiding principles, intended to make biological data discovery and use easier for both humans and machines.

Findable – Data should have unique and persistent identifiers
Accessible – Authorized users can retrieve it through standard protocols
Interoperable – Shared vocabularies and machine-readable metadata make data usable across systems
Reusable – Clear usage licenses and detailed descriptions allow future reuse

FAIR principles were developed in response to the field’s reproducibility crises and fragmented data infrastructures. They align with traditional research ethics of transparency, accountability, and reproducibility. And, in the context of AI, FAIR data allow model outcomes to be audited and traced. 

These are not just technical standards. They are ethical imperatives, essential to underpin trust in AI-driven biological research. 

And data alone is not enough. We also need rich and accurate metadata, fully attached to data points and findable by the same unique identifier.

Our newsletter bridges the gap between AI engineering, philosophy, and public debate. Sign up now.

Another example: a researcher is measuring a protective chemical produced by maple leaves under attack. She processes her samples, runs the data, and notices one leaf is an outlier. The data point itself only tells her that this leaf shows a much higher concentration than the other leaves. But — when she refers to her notes — she sees the leaf was lunch for a caterpillar.

Metadata, “the data about the data” — how the sample was taken, what time of day, how it was stored, any distinguishing features — is crucial information. It can tell you why a data point is where it is. It is vital to scientific understanding. 

Without rich metadata AI cannot fully interpret or reproduce results, risking false conclusions and biased models. 

There is a growing ethical tension between the acceleration of knowledge and humanity’s capacity to control it. FAIR data and ethical governance will be essential to ensure AGI augments rather than endangers biological progress.

AI already offers transformative benefits for biology, but its reliability depends on FAIR, ethically managed data. And AGI could revolutionise discovery — but only if it is grounded in transparent, accountable systems.

The transition from AI to AGI in life sciences should be guided by human values codified in ethical data practice. Investing in FAIR infrastructure, open standards, and ethical oversight is essential to ensure that AGI enhances rather than endangers biological progress.

We need collaboration between researchers, data professionals, and policymakers to create responsible progress towards AGI in biology. Ethical principles and clean, FAIR data standards must be embedded into every stage of research. This demands a robust framework:

Accountability: who is responsible for errors, bias, or misuse in AI and AGI-generated biological outputs?

Transparency: explainability and maintained audit trails for AI-driven experimentation.

Beneficence and non-maleficence: avoiding harm in biotechnology applications.

Transparent data and metadata are the ethical infrastructure ensuring oversight and traceability. Shared norms for biological data and model exchange must be evolved and ethicists and civil society need to be included in decision-making.

Ethics and FAIR data governance must co-evolve with AI sophistication and AGI development should never outpace the systems ensuring safety and accountability.

Responsible data practice is not just an administrative task. It’s the foundation of an ethical, innovative, and secure biological future.

Want the latest AGI news and research summaries delivered to your inbox every month? Subscribe now to AGI Ethics News.

Did you enjoy this article? Share it with a friend!