Research area
Machine learning for biomaterials discovery
How high-throughput experimentation and machine learning map biomaterial structure-function behaviour: methods, dataset sizes and descriptors.
Biomaterials are difficult to design because performance rarely traces to a single property. Surface chemistry, molecular weight distribution, charge density, hydrophobic balance and architecture all contribute, and they interact. Holding everything constant and varying one thing samples a line through a space that has structure in every direction.
That is the case for a different method, and it is the argument my review in Tissue Engineering Part A sets out: pair high-throughput experimentation with machine learning, and map the structure–function surface rather than probing it point by point.
How big is a high-throughput biomaterials dataset?
Study sizes cited in the review, on a logarithmic axis. Four orders of magnitude separate a careful polymer library from the combinatorial space it samples.
The gap is the argument. Experimental libraries reach a few thousand members; the spaces they are drawn from run to millions. Nothing exhaustive is possible here, so the question becomes which few thousand to make, which is a modelling problem, not a pipetting one.
Show the numbers
| Study | Size | Description |
|---|---|---|
| Kohn polyacrylate library | 112 | 112 distinct degradable polymers |
| Scaffold study | 182 | 13 polymers, 182 scaffolds tested |
| Titanium nanotube analysis | 272 | 272 labelled samples, 30 publications |
| Lipid nanoparticle screen | 1,080 | 1,080-LNP formulation library |
| Poly(β-amino ester) library | 2,500 | ~2,500 degradable PAEs for gene delivery |
| Drug–excipient space | 2,100,000 | 2.1 million possible pairings |
Reported Study sizes as cited in Ahmed et al., Tissue Engineering Part A 30(19–20), 662–680 (2024). A dot plot is used rather than bars because a bar length on a log axis misrepresents ratio.
The methods, and where they are actually used
There is no single algorithm for biomaterials. What gets used depends on how much data exists, whether the target is continuous or categorical, and whether the point is prediction or working out which features matter. Random forests earn their place partly because feature importance falls out of them; Gaussian regression suits small datasets with useful uncertainty estimates; active learning fits the design–build–test–learn loop that automated synthesis makes possible.
Which methods are being used where
The pairings of algorithm and application area discussed in the review.
Show the numbers
| Method | Application areas discussed |
|---|---|
| Random forest | Tissue engineering, Antifouling materials, Data mining |
| Support vector machine | Tissue engineering, Antifouling materials |
| Neural network | Gene delivery, Drug delivery |
| Gradient boosting | Gene delivery |
| Gaussian regression | Tissue engineering, Drug delivery |
| Hidden Markov model | Protein stabilization |
| Active learning | Drug delivery, Protein stabilization |
Reported Pairings as discussed in the review. A filled dot means the review covers that method in that area; an empty position is not a claim that the combination is unused.
A model only sees the descriptors you chose
This is the step that decides what is learnable. A model has no access to a polymer; it has access to the numbers used to represent it. Choose descriptors that miss the property driving behaviour and no amount of data or model capacity recovers it.
What a polymer looks like to a model
A model cannot read a structure. It reads whatever numbers you chose to describe it, and that choice bounds what the model can possibly learn.
Show the numbers
| Descriptor family | Examples |
|---|---|
| Structural | Fibre diameter, pore diameter, porosity, pore size |
| Mechanical | Young's modulus, compressive strength |
| Thermal | Glass transition temperature (Tₘ) |
| Surface | Contact angle, roughness, hydrophobicity, wettability |
| Compositional | Monomer ratios, pendant groups, backbone modifications |
| Molecular | Cheminformatic descriptors (e.g. PaDEL) for lipids and small molecules |
Reported Descriptor families as surveyed in the review.
The literature is the wrong training set
Published biomaterials results are a filtered sample. Papers report formulations that worked. " "The ones that aggregated, failed to release, or provoked a response are largely absent, not " "through dishonesty, but because null results are hard to publish and the material was dropped.
A model trained on that record learns which successful materials resemble other successful materials. It has little to say about where the useful region ends, because it has never been shown the other side of the boundary. High-throughput data is different in kind: when you run a plate, you keep every well. The formulations that precipitated are recorded with the same rigour as the ones that performed, because the same instrument measured them in the same run.
The paper behind this page
Mapping Biomaterial Complexity by Machine Learning
Ahmed, Mulay, Ramirez et al., Tissue Engineering Part A, 2024.
doi:10.1089/ten.tea.2024.0067