A credible modeled interaction does not imply a developable protein — and a developable sequence does not imply a credible binder.
In this campaign, developability means the practical sequence-level properties that determine whether a designed binder can plausibly move from a computational model into an experimental workflow. A good interface model is not enough; the isolated binder also has to look expressible, soluble, stable, and not obviously aggregation-prone.
Developability does not ask whether the binder engages PD-L1. It asks whether the binder sequence looks like a workable protein before experimental testing.
Expression, solubility, thermostability, aggregation propensity, and folding stability can improve or worsen independently of interface geometry.
The goal is not to declare winners from sequence prediction alone. The goal is to avoid over-prioritizing candidates that only look good structurally.
Because these are de novo mini-binders, Prot2Prop outputs are best treated as relative ranking signals, not absolute forecasts.
The structural and MD pipeline from the previous stage of this campaign asks whether a designed complex is geometrically plausible and dynamically stable. But a protein that looks good in a simulation still needs to be produced, purified, and handled in a laboratory. Expression yield, solubility, aggregation propensity, and thermal stability are distinct from interface quality — and often in tension with it. All 55 designed binder sequences were screened through Prot2Prop to ask whether sequence-level developability predictions add information beyond the structural pipeline.
The overall picture is encouraging: most designs are predicted to be producible and soluble, with moderate thermostability. Expression yield and folding stability scores skew negative, but these are relative scales — without calibration data for de novo mini-binders, the absolute values are less meaningful than the ranking.
The central finding: almost every correlation between Prot2Prop metrics and pipeline metrics is non-significant. Prot2Prop is measuring something the structural pipeline does not capture.
Two correlations stand out. Prot2Prop solubility correlates modestly with ESMFold pLDDT (ρ = +0.386, p = 0.004), which makes physical sense — sequences predicted to fold confidently are also predicted to be soluble. Both respond to whether the sequence "looks like a well-behaved protein."
The most provocative finding: designs with more interfacial contacts are predicted to be less producible (ρ = −0.530, p = 0.0037). This negative correlation suggests an experimentally useful hypothesis: higher contact counts may be reporting interface size, surface composition, binder length, or a Prot2Prop calibration effect on de novo designs.
Prot2Prop predictions show a clear dependence on binder length that maps to physical intuition.
The 25-residue designs show near-perfect predicted solubility (1.000) but poor folding stability (−1.43). The 100-residue designs have the best folding stability (−0.97) but somewhat lower solubility (0.86). Short peptides are soluble but may lack stable tertiary structure, while longer proteins fold more stably — a solubility–stability tradeoff that is a familiar theme in protein engineering. (Aggregation propensity was not stratified by length in this screen, so it isn't part of this comparison.)
Adding six Prot2Prop columns to the retrospective pipeline matrix changes the composite ranking. The top group is broadly stable — the same general set of designs remains competitive — but the internal ordering shifts, and several previously low-ranked designs rise sharply.
| Design | Hotspot | Part 7 rank | Augmented rank | Δ Rank | Part 7 composite | Augmented composite |
|---|
| Rank | Design | Hotspot | Augmented composite | Part 7 composite |
|---|
The augmented top 10 is more diverse than the Part 7 top 10: distributed designs hold ranks 2–3, clusterB appears at ranks 5–7, and a previously invisible clusterA design enters at rank 10. This broadening of the candidate set is a practical benefit of adding developability screening.
An important methodological caution. This screen adds six developability columns to a matrix that previously contained fourteen pipeline metrics (see Part 7). For designs with all 20 rank-normalized columns available, developability contributes 6/20 = 30% of the augmented composite — not because anyone decided developability is exactly 30% as important as the rest of the pipeline, but because it contributes many columns. For designs missing downstream metrics, the effective developability weight can be even higher because the composite is averaged over available columns.
A more principled approach would weight by evidence family — giving each of the seven families (fold quality, complex prediction, interface geometry, affinity proxy, MD stability, contact persistence, developability) equal voice regardless of how many individual metrics each contains. The pipeline does not implement this family-weighted composite, which is itself a finding: the architecture of the scoring system matters as much as the individual scores.
Rather than collapsing every metric into one composite number, a more defensible next step treats designability and developability as separate, only loosely correlated axes and applies explicit judgment at their intersection — including a few designs chosen deliberately as longshots or negative controls, so the predictions themselves get tested alongside the candidates.
The six Prot2Prop outputs are not fully redundant, but they are also not independent. Several modest internal correlations appear, including material production with solubility (ρ = −0.320, p = 0.017), temperature stability (ρ = −0.309, p = 0.022), and expression yield (ρ = +0.280, p = 0.039); temperature stability with folding stability (ρ = +0.454, p = 0.0005); and aggregation propensity with folding stability (ρ = −0.335, p = 0.013). This supports treating Prot2Prop as an evidence family with internal structure, not as six completely independent votes.