Blog
← All posts

Data · AI in life sciences

Dark data: science below the waterline

Life sciences does not lack data. It lacks usable records of what happened at the bench—and AI cannot learn from evidence it cannot see.

The scientific literature is the visible tip of an iceberg.

Above the waterline sit papers, curated datasets and successful results: searchable, citable and ready to train models. Below it sit the exploratory runs, failed experiments, protocol changes, handwritten observations and instrument exports that made those results possible.

This is dark data. It is not necessarily secret, lost or poor quality. It is simply hard to find, interpret or reuse—even inside the lab that generated it.

Life sciences does not lack data. It lacks data with enough structure, provenance and context to become reliable evidence.

The UK government’s AI Adoption Plan for Life Sciences puts the problem plainly: biological data is fragmented, historical records demand expensive wrangling, and unpublished or non-machine-readable results hold enormous untapped value.

If AI is to transform life sciences, we first have to illuminate what is below the waterline.

What makes scientific data dark?

Dark data is more than an unpublished spreadsheet. A record becomes dark whenever the people or machines that could learn from it cannot discover, understand or use it.

In a typical research workflow, this can include:

  • a failed assay recorded only in a notebook;
  • a plate-reader export with an ambiguous filename;
  • a result detached from the protocol version that produced it;
  • a method change buried in a message;
  • an image without a sample ID or acquisition conditions;
  • a useful negative result that was never published; or
  • data with unclear consent, ownership or permitted uses.

These are different problems, but they have the same outcome: the experiment happened, resources were spent and knowledge was created, yet that knowledge cannot travel.

More files do not solve the problem. A folder can be full and still be dark. A useful result must stay connected to its question, method, inputs, conditions, observations and authorship.

Why does so much data stay below the surface?

Science is rewarded at the point of publication

The scientific record selects for findings novel, positive and complete enough to become papers. A meta-analysis of publication bias found that studies with positive or significant results were more likely to be published. A systematic review of unpublished medical research found barriers ranging from negative results to fear of rejection and authorship problems. Useful failures rarely cross the publication threshold.

A paper is also a summary. It communicates a defensible conclusion, not every exploratory run, troubleshooting note or discarded parameter. That compression helps readers, but hides much of the experimental history.

The research workflow is fragmented

The method is a PDF. The plan is in a notebook. Execution lives on a paper checklist, raw output in an instrument file and interpretation in a slide deck. Messages fill the gaps. Each tool does its job, but a person must remember how the pieces connect. A result is only as useful as its experimental context.

When that person changes project or leaves the lab, the files may remain while their meaning disappears.

Context is treated as paperwork

Data stewardship often waits until the experiment is complete. By then, vital details have faded: which protocol was open, why an incubation ran long, which sample looked unusual. Reconstruction is expensive and unreliable. Double entry makes it less likely to happen at all.

Not all data should be open

Life-science records can contain intellectual property, commercially sensitive findings and health information. Visible does not mean public. Access must reflect consent, ownership and legitimate use. The NIH’s data-sharing guidance makes the same distinction.

The real goal is usable under appropriate governance: discoverable to the right people, with clear permissions and provenance.

Why dark data is an AI problem

An AI model can only learn from evidence it can access and interpret. If its record over-represents polished successes, it sees a distorted version of science. It learns what worked without seeing the boundary conditions, failed routes and practical decisions behind that success.

Negative results define the edge of the search space. A well-recorded failure can rule out a target, concentration, material or condition. Without it, another scientist—or an automated lab—may explore the same dead end.

Context also separates a useful training example from a misleading correlation. A measurement without its method, units, sample provenance and conditions may add volume to a dataset while reducing its quality. More data is not the same as more evidence.

That is why the FAIR Guiding Principles—Findable, Accessible, Interoperable and Reusable—emphasise machine actionability and provenance. AI readiness begins before a model is chosen. It begins when the experiment is recorded.

How we bring dark data into the light

No single database will solve this. Change must begin at the bench.

Capture structure as the work happens

Context is cheapest at the moment of experimentation. The protocol, samples, parameters and researcher are already known. Software should preserve those relationships automatically and ask for input only when it matters—not years later, during an archive clean-up.

Record execution, not just intention

A protocol says what was meant to happen. An experimental record must preserve what did happen. Deviations, failures and observations belong beside the method, not tidied away.

That is also why we treat protocols as versioned building blocks, rather than static instructions copied into each new experiment.

Treat negative results as knowledge

Every experiment deserves a durable place in the research history. A result need not become a paper to prevent duplication, improve a protocol or shape the next design.

Make records machine-actionable without making scientists into data engineers

Structured fields, explicit units, versioned methods and linked provenance make records easier for software to compare. But the interface must fit the practical, iterative nature of wet-lab work. Too much friction sends record-keeping straight back to paper.

Separate visibility from access

Good governance makes clear what exists, who can use it and why. Sensitive data can remain protected yet reusable inside an authorised team or trusted research environment. Sharing should be deliberate, not a side effect of making data useful.

Sous.bio’s approach: make every experiment count

Sous is being built around a simple idea: the experiment—not the document—is the unit of scientific knowledge.

We reconnect the parts of experimental work that usually drift apart:

  • Protocols are structured and versioned. Every experiment can remain linked to the exact method used, even after the shared protocol evolves.
  • The run preserves reality. Notes, deviations and observations describe what happened at the bench without silently changing what was planned.
  • Results keep their context. Files and conclusions stay connected to the experiment, its samples and its place in the wider project.
  • Research forms a history. Related experiments can be linked from an early pilot through validation to follow-up work—including the attempts that failed.
  • Capture fits the scientist. Researchers can work through a protocol at the bench or bring handwritten notes into the record, reducing the need to reconstruct the experiment later.
  • AI assists by choice. Sous can use published literature alongside a team’s own experimental history to help critique plans and conclusions, while keeping the scientist responsible for the decision.

This does not require a lab to publish unfinished work. Raw data and unpublished results stay private until the team chooses otherwise, and they are not used to train models. The first benefit is local: colleagues can understand old results, learn from failed runs and plan from more than collective memory.

Over time, connected records create something larger. With the right permissions, data can be compared across runs, shared responsibly and used by computational tools—without months of forensic cleaning.

The opportunity beneath the waterline

Published science will always be selective. It should be. Papers communicate conclusions, not every moment of the research process.

But evidence that never reaches a paper should not disappear. The setup that explains a result, the observation that changes a protocol and the failed experiment that closes a dead end all have value.

The next leap in life-science AI will not come from larger models and more compute alone. It will come from treating everyday experimental work as durable, connected and governable knowledge.

The tip of the iceberg has already taught us a great deal. Imagine what becomes possible when we can learn responsibly from the rest.

Further reading