Eman AhmedRutgers BME

Research area

Machine learning for biomaterials discovery

How high-throughput experimentation and machine learning map biomaterial structure-function behaviour: methods, dataset sizes and descriptors.

Biomaterials are difficult to design because performance rarely traces to a single property. Surface chemistry, molecular weight distribution, charge density, hydrophobic balance and architecture all contribute, and they interact. Holding everything constant and varying one thing samples a line through a space that has structure in every direction.

That is the case for a different method, and it is the argument my review in Tissue Engineering Part A sets out: pair high-throughput experimentation with machine learning, and map the structure–function surface rather than probing it point by point.

How big is a high-throughput biomaterials dataset?

Study sizes cited in the review, on a logarithmic axis. Four orders of magnitude separate a careful polymer library from the combinatorial space it samples.

Dot plot on a log axis. Scaffold study 182, titanium nanotube analysis 272, Kohn polyacrylate library 112, lipid nanoparticle screen 1080, poly(beta-amino ester) library 2500, drug-excipient space 2.1 million.10²10³10⁴10⁵10⁶10⁷Kohn polyacrylate library: 112Kohn polyacrylate library112 distinct degradable polymers112Scaffold study: 182Scaffold study13 polymers, 182 scaffolds tested182Titanium nanotube analysis: 272Titanium nanotube analysis272 labelled samples, 30 publications272Lipid nanoparticle screen: 1,080Lipid nanoparticle screen1,080-LNP formulation library1,080Poly(β-amino ester) library: 2,500Poly(β-amino ester) library~2,500 degradable PAEs for gene delivery2,500Drug–excipient space: 2,100,000Drug–excipient space2.1 million possible pairings2,100,000Number of distinct formulations or samples (log scale)

The gap is the argument. Experimental libraries reach a few thousand members; the spaces they are drawn from run to millions. Nothing exhaustive is possible here, so the question becomes which few thousand to make, which is a modelling problem, not a pipetting one.

Show the numbers
StudySizeDescription
Kohn polyacrylate library112112 distinct degradable polymers
Scaffold study18213 polymers, 182 scaffolds tested
Titanium nanotube analysis272272 labelled samples, 30 publications
Lipid nanoparticle screen1,0801,080-LNP formulation library
Poly(β-amino ester) library2,500~2,500 degradable PAEs for gene delivery
Drug–excipient space2,100,0002.1 million possible pairings

Reported Study sizes as cited in Ahmed et al., Tissue Engineering Part A 30(19–20), 662–680 (2024). A dot plot is used rather than bars because a bar length on a log axis misrepresents ratio.

The methods, and where they are actually used

There is no single algorithm for biomaterials. What gets used depends on how much data exists, whether the target is continuous or categorical, and whether the point is prediction or working out which features matter. Random forests earn their place partly because feature importance falls out of them; Gaussian regression suits small datasets with useful uncertainty estimates; active learning fits the design–build–test–learn loop that automated synthesis makes possible.

Which methods are being used where

The pairings of algorithm and application area discussed in the review.

Dot matrix pairing seven algorithms with six application areas.Tissue engineeringGene deliveryDrug deliveryProtein stabilizationAntifouling materialsData miningRandom forestRandom forest · Tissue engineeringRandom forest · Antifouling materialsRandom forest · Data miningSupport vector machineSupport vector machine · Tissue engineeringSupport vector machine · Antifouling materialsNeural networkNeural network · Gene deliveryNeural network · Drug deliveryGradient boostingGradient boosting · Gene deliveryGaussian regressionGaussian regression · Tissue engineeringGaussian regression · Drug deliveryHidden Markov modelHidden Markov model · Protein stabilizationActive learningActive learning · Drug deliveryActive learning · Protein stabilization
Show the numbers
MethodApplication areas discussed
Random forestTissue engineering, Antifouling materials, Data mining
Support vector machineTissue engineering, Antifouling materials
Neural networkGene delivery, Drug delivery
Gradient boostingGene delivery
Gaussian regressionTissue engineering, Drug delivery
Hidden Markov modelProtein stabilization
Active learningDrug delivery, Protein stabilization

Reported Pairings as discussed in the review. A filled dot means the review covers that method in that area; an empty position is not a claim that the combination is unused.

A model only sees the descriptors you chose

This is the step that decides what is learnable. A model has no access to a polymer; it has access to the numbers used to represent it. Choose descriptors that miss the property driving behaviour and no amount of data or model capacity recovers it.

What a polymer looks like to a model

A model cannot read a structure. It reads whatever numbers you chose to describe it, and that choice bounds what the model can possibly learn.

Show the numbers
Descriptor familyExamples
StructuralFibre diameter, pore diameter, porosity, pore size
MechanicalYoung's modulus, compressive strength
ThermalGlass transition temperature (Tₘ)
SurfaceContact angle, roughness, hydrophobicity, wettability
CompositionalMonomer ratios, pendant groups, backbone modifications
MolecularCheminformatic descriptors (e.g. PaDEL) for lipids and small molecules

Reported Descriptor families as surveyed in the review.

The literature is the wrong training set

Published biomaterials results are a filtered sample. Papers report formulations that worked. " "The ones that aggregated, failed to release, or provoked a response are largely absent, not " "through dishonesty, but because null results are hard to publish and the material was dropped.

A model trained on that record learns which successful materials resemble other successful materials. It has little to say about where the useful region ends, because it has never been shown the other side of the boundary. High-throughput data is different in kind: when you run a plate, you keep every well. The formulations that precipitated are recorded with the same rigour as the ones that performed, because the same instrument measured them in the same run.

More on this in the research notes.

The paper behind this page

Mapping Biomaterial Complexity by Machine Learning
Ahmed, Mulay, Ramirez et al., Tissue Engineering Part A, 2024. doi:10.1089/ten.tea.2024.0067

Other research areas