Eman AhmedRutgers BME

Research note · 6 min read

Biomaterials discovery needs the experiments that failed

By Eman Ahmed, PhD candidate, Gormley Lab, Rutgers University

Biomaterials are hard to design because their performance usually comes from several subtle properties interacting at once. Surface chemistry, molecular weight distribution, charge density, hydrophobic balance and architecture all contribute, and rarely independently. The structure–function relationship is real, but it is not something you can usually write down.

The traditional response is to hold everything constant and vary one thing. It is rigorous, and in a space with this much interaction between variables it is close to hopeless. You sample a line through a space that has structure in every direction.

The obvious fix has a less obvious problem

Machine learning is the natural tool for a high-dimensional structure–function map, and the barrier to using it has dropped sharply: an experimentalist can now train a competent model without a background in statistics. That is a genuine shift, and it is the shift our review in Tissue Engineering Part A was written around.

But models are only as good as what they learn from, and the biomaterials literature is a filtered sample. Papers report the formulations that worked. The ones that aggregated, or failed to release, or provoked an immune response, are mostly absent, not through dishonesty, but because a null result is hard to publish and the material was abandoned.

A model trained on that record learns which successful materials resemble other successful materials. It has very little to say about where the boundary of the useful region lies, because it has never been shown the other side of it.

Why high-throughput data is different in kind

This is the part I think is underappreciated. When you run a 96-well plate, you keep every well. The formulations that precipitated are recorded with the same rigour as the ones that performed, because they were measured by the same instrument in the same run. You are not making a publication decision about each data point.

The resulting dataset is balanced in a way that a literature-derived one structurally cannot be. That is a large part of why high-throughput experimentation and machine learning belong together, not because the throughput is impressive, but because of what it does to the shape of the data.

Where data mining still helps

None of which means the published record is useless. Text and data mining across existing literature can map regions of a design space cheaply enough to tell you where not to spend benchtime, and for well-studied material classes the aggregate signal is strong. Our review covers these approaches alongside direct experimentation, across tissue engineering, gene delivery, drug delivery, protein stabilization and antifouling materials.

The practical position is that mined data is good for narrowing and high-throughput data is good for deciding. Used the other way round, you get a model that is confident about a region nobody has actually measured.

What follows for how we run experiments

If you accept that the failures carry information, some things follow. Record them with the same metadata as the successes. Do not discard a plate because most of it did not work. Report the full screened range, not the subset that supports the conclusion. And design assays so that failure produces a number rather than an absence: a solubility of zero is data; a well you stopped measuring is not.

Source

This note summarises Mapping Biomaterial Complexity by Machine Learning (Ahmed, Mulay, Ramirez et al., Tissue Engineering Part A, 2024; doi:10.1089/ten.tea.2024.0067). These notes are plain-language companions to peer-reviewed work. Every factual claim traces to the paper linked at the end of the note.

  • Machine learning
  • Biomaterials
  • Research practice