← Back to project hub

Sequence-based developability screening

A credible modeled interaction does not imply a developable protein — and a developable sequence does not imply a credible binder.

Prot2Prop ProstT5 Sequence-only prediction Solubility Expression Thermostability Aggregation
Prot2Prop (Gharaie Amirabadi et al., 2026, bioRxiv) uses a frozen ProstT5 encoder with shared and task-specific residual adapters to predict six developability properties from amino acid sequence input. It was trained on natural proteins — performance on de novo mini-binders has not been benchmarked. Predictions are for monomer properties, not binding affinity or complex stability. This is a preprint, not yet peer-reviewed.

What is developability?

In this campaign, developability means the practical sequence-level properties that determine whether a designed binder can plausibly move from a computational model into an experimental workflow. A good interface model is not enough; the isolated binder also has to look expressible, soluble, stable, and not obviously aggregation-prone.

Not a binding score

Developability does not ask whether the binder engages PD-L1. It asks whether the binder sequence looks like a workable protein before experimental testing.

Orthogonal evidence

Expression, solubility, thermostability, aggregation propensity, and folding stability can improve or worsen independently of interface geometry.

Useful for triage

The goal is not to declare winners from sequence prediction alone. The goal is to avoid over-prioritizing candidates that only look good structurally.

Model-calibration dependent

Because these are de novo mini-binders, Prot2Prop outputs are best treated as relative ranking signals, not absolute forecasts.

Why developability matters

The structural and MD pipeline from the previous stage of this campaign asks whether a designed complex is geometrically plausible and dynamically stable. But a protein that looks good in a simulation still needs to be produced, purified, and handled in a laboratory. Expression yield, solubility, aggregation propensity, and thermal stability are distinct from interface quality — and often in tension with it. All 55 designed binder sequences were screened through Prot2Prop to ask whether sequence-level developability predictions add information beyond the structural pipeline.

Prot2Prop predictions across 55 designs

Material production
Classification
73% predicted producible
Solubility
Classification
91% predicted soluble
Thermostability
Classification
67% predicted thermostable
Aggregation propensity
Regression
6.1 ± 3.2 (lower = better)
Expression yield
Regression
−0.59 ± 1.01
Folding stability
Regression
−1.23 ± 0.70

The overall picture is encouraging: most designs are predicted to be producible and soluble, with moderate thermostability. Expression yield and folding stability scores skew negative, but these are relative scales — without calibration data for de novo mini-binders, the absolute values are less meaningful than the ranking.

Developability is mostly independent of the structural pipeline

The central finding: almost every correlation between Prot2Prop metrics and pipeline metrics is non-significant. Prot2Prop is measuring something the structural pipeline does not capture.

Cross-stage correlations: Prot2Prop vs. pipeline metrics (Spearman ρ)

Two correlations stand out. Prot2Prop solubility correlates modestly with ESMFold pLDDT (ρ = +0.386, p = 0.004), which makes physical sense — sequences predicted to fold confidently are also predicted to be soluble. Both respond to whether the sequence "looks like a well-behaved protein."

The interface quality–producibility tension

The most provocative finding: designs with more interfacial contacts are predicted to be less producible (ρ = −0.530, p = 0.0037). This negative correlation suggests an experimentally useful hypothesis: higher contact counts may be reporting interface size, surface composition, binder length, or a Prot2Prop calibration effect on de novo designs.

ρ = −0.530
Spearman correlation
Material-production probability versus static interface contact count.
p = 0.0037
Nominal significance
Computed across designs with static-interface metrics available.
n = 28
Designs compared
The 28 of 55 designs that reached full downstream characterization in Part 7.
This anticorrelation is the most experimentally actionable hypothesis from the developability screen. If confirmed, the pipeline faces a real tradeoff between interface quality and monomer manufacturability — and construct selection needs to balance both, not simply maximize contact scores. If not confirmed, it still flags where model calibration on de novo designs needs to be tested experimentally.

Length-dependent patterns

Prot2Prop predictions show a clear dependence on binder length that maps to physical intuition.

Developability predictions by binder length

The 25-residue designs show near-perfect predicted solubility (1.000) but poor folding stability (−1.43). The 100-residue designs have the best folding stability (−0.97) but somewhat lower solubility (0.86). Short peptides are soluble but may lack stable tertiary structure, while longer proteins fold more stably — a solubility–stability tradeoff that is a familiar theme in protein engineering. (Aggregation propensity was not stratified by length in this screen, so it isn't part of this comparison.)

Ranking shifts when developability is added

Adding six Prot2Prop columns to the retrospective pipeline matrix changes the composite ranking. The top group is broadly stable — the same general set of designs remains competitive — but the internal ordering shifts, and several previously low-ranked designs rise sharply.

DesignHotspotPart 7 rankAugmented rankΔ RankPart 7 compositeAugmented composite
The biggest mover — len100_clusterA_noise0__design_0_0_rank7 — jumps 40 positions, from pipeline rank 50 to augmented rank 10. This design scored poorly on structural and MD metrics but has favorable predicted developability. Whether that translates to a better experimental candidate is unknowable without data, but it is exactly the kind of re-ranking that an orthogonal evidence layer should produce. If all it did was confirm the existing ranking, it would be redundant.

Augmented top 10

RankDesignHotspotAugmented compositePart 7 composite

The augmented top 10 is more diverse than the Part 7 top 10: distributed designs hold ranks 2–3, clusterB appears at ranks 5–7, and a previously invisible clusterA design enters at rank 10. This broadening of the candidate set is a practical benefit of adding developability screening.

Equal-weight composites contain hidden choices

An important methodological caution. This screen adds six developability columns to a matrix that previously contained fourteen pipeline metrics (see Part 7). For designs with all 20 rank-normalized columns available, developability contributes 6/20 = 30% of the augmented composite — not because anyone decided developability is exactly 30% as important as the rest of the pipeline, but because it contributes many columns. For designs missing downstream metrics, the effective developability weight can be even higher because the composite is averaged over available columns.

A more principled approach would weight by evidence family — giving each of the seven families (fold quality, complex prediction, interface geometry, affinity proxy, MD stability, contact persistence, developability) equal voice regardless of how many individual metrics each contains. The pipeline does not implement this family-weighted composite, which is itself a finding: the architecture of the scoring system matters as much as the individual scores.

Rather than collapsing every metric into one composite number, a more defensible next step treats designability and developability as separate, only loosely correlated axes and applies explicit judgment at their intersection — including a few designs chosen deliberately as longshots or negative controls, so the predictions themselves get tested alongside the candidates.

Prot2Prop internal structure

The six Prot2Prop outputs are not fully redundant, but they are also not independent. Several modest internal correlations appear, including material production with solubility (ρ = −0.320, p = 0.017), temperature stability (ρ = −0.309, p = 0.022), and expression yield (ρ = +0.280, p = 0.039); temperature stability with folding stability (ρ = +0.454, p = 0.0005); and aggregation propensity with folding stability (ρ = −0.335, p = 0.013). This supports treating Prot2Prop as an evidence family with internal structure, not as six completely independent votes.

What this means

Prot2Prop predictions are largely orthogonal to the structural and MD pipeline, with two notable exceptions: solubility tracks fold confidence, and predicted producibility falls as interfacial contacts increase. Developability forms a separate evidence layer that addresses a separate question: not "does this look like a plausible binder?" but "does this sequence look like a producible protein?" Designs that score well across both the structural pipeline and the developability screen are stronger experimental candidates because they have survived largely independent challenges. The production-vs-contacts anticorrelation is a genuine design tension that should shape construct selection — not a problem to be optimized away.