Designed Proteins Are Leaving the Lab. The Validation Gap Is Widening.
Computational design now produces candidate binders in days. Experimental characterisation still takes months, and the backlog is changing which claims get published.

The design half of protein engineering has become extraordinarily fast. A group with a target structure and reasonable compute can generate thousands of candidate binders in a week, and a meaningful fraction of them will fold as predicted. The characterisation half has not sped up at all. Expressing, purifying, and measuring a candidate still takes weeks per construct, and the good measurements take longer than that.
That asymmetry is not a temporary inconvenience. It is starting to shape what the field publishes and what it believes.
Where the ratio actually sits
Groups we spoke to described design-to-validation ratios that would have been unthinkable five years ago. One lab generated roughly 4,000 candidate designs against a single target over two months and experimentally characterised 47 of them.
The 47 were not a random sample. They were selected by computational filters, which is reasonable practice and also the source of the problem. The published result describes the performance of the selected subset. It does not describe the performance of the design method, because the selection step is doing work that is not being measured.
- Filtering is a hidden model. Predicted binding affinity, predicted solubility, and structural plausibility scores all encode assumptions. Reporting success only among survivors measures the filters as much as the generator.
- Negative results stay unpublished. A design campaign that yields nothing is rarely written up, so the field’s estimate of base rates is drawn from campaigns that worked.
- Characterisation depth varies enormously. A binding assay is not a functional assay, and a functional assay in vitro is not activity in a cell.
We can now design faster than we can be wrong at a measurable rate. That should worry people more than it does.
The measurements that are being skipped
The specific gap most often cited by structural biologists is not affinity. It is specificity and stability under realistic conditions.
Affinity for the intended target is comparatively easy to measure and is almost always reported. Off-target binding across a realistic proteome is expensive and is usually not. Thermal and proteolytic stability get reported inconsistently. Aggregation behaviour at concentration, which determines whether a molecule is developable at all, appears in a minority of papers.
This produces a literature in which designs look excellent on the axis that is cheap to measure. Groups working on therapeutic applications are blunt about the consequence: a substantial share of published designed binders fail on properties that were never characterised in the original report.
What would close the gap
Three interventions came up repeatedly, and none of them requires a methodological breakthrough.
The first is reporting the denominator. State how many designs were generated, what filters were applied, and how many survived each stage. This is a change in convention rather than in capability, and it would immediately make published success rates interpretable.
The second is standardised minimum characterisation. A short, agreed panel covering specificity, stability, and aggregation, reported for every candidate that gets published, would eliminate most of the current inconsistency. Several groups are pushing for this through journal policy rather than waiting for consensus.
The third is investment in throughput on the wet side. Automated expression and purification exists and works. It is unglamorous, it does not produce papers on its own, and it is chronically underfunded relative to the compute budgets on the design side.
Why this is a familiar failure
The pattern here is not specific to protein design. It is the standard signature of a field where one half of the loop got cheap and the other did not: apparent progress accelerates, published success rates rise, and the base rate quietly becomes unknowable.
Medicine has run this experiment already, which is why trial registration exists at all. As we reported on the trial reporting gap, even mandatory registration only partly solved it. The design field has the advantage of being able to adopt the convention before the credibility problem becomes acute rather than after.
Published . Corrections and clarifications: our policy.


