Building protein models that understand biophysics

Building protein models that understand biophysics

Almost every major advance in applying AI to biology is preceded by data collection and curation. At Escalante, we're building scalable assays to enable the next generation of models.

Protein models are remarkably good at predicting structures and designing binders, but in many applications this is not enough: useful molecules must bind the right targets, with the right kinetics, while avoiding consequential off-targets (and satisfying a host of other constraints). Current models are not calibrated to predict these properties across arbitrary protein pairs, largely because suitable training data do not exist.

We argue that a major missing ingredient is large, standardized protein-protein interaction data containing hard negatives and positives, sparse library-on-library measurements, dense mutational scans, and separate kon and koff kinetics measurements across many binders and targets. To address this, we're developing a pooled library-on-library assay designed to measure on the order of one million pairwise interactions per run. In today's post, we'll describe the motivation and share the initial results showing we can measure kinetics from DNA sequencing reads alone.

Do we need to improve structure prediction and design?

For most binder design applications, generating a binder that hits its intended target is not enough. For instance, in a therapeutic program, the binder must also have suitable association and dissociation rates, avoid consequential off-target interactions, and satisfy many other constraints[1][2][[1]]. Ideally, more of these constraints would be predicted and considered during in silico design, rather than discovered through expensive wet-lab tests — or even further downstream, when they can produce program-killing side effects.

Binding specificity[[2]] illustrates how demanding some of these constraints can be for models. Unexpected cross-reactivity can cause severe toxicity, as in the well-known case of engineered MAGE-A3-directed T cells recognizing a titin-derived peptide[3]. As a rough[[3]] thought experiment, suppose we want to do in silico off-target screening against roughly 20,000 candidate human proteoforms. Even a false-positive rate of 10−3 — far, far better than current general-purpose cofolding benchmarks demonstrate — would produce about 20 false positives. Averaging fewer than one false positive would require achieving a false-positive rate below approximately 5 × 10−5, while still maintaining a reasonable true-positive rate.

The required accuracy is therefore very different from the accuracy needed to enrich for binders in the first stage of a binder design campaign. Likewise, affinity alone is an incomplete target: two interactions with the same KD can have different kon and koff, and therefore different biological behavior[4].

What are the limits of structure prediction?

Current protein models are strikingly effective at structure prediction, generation, and docking of interaction partners. But, for the most part, they cannot predict whether arbitrary proteins bind, how tightly they bind, or with what kinetics and specificity.

A useful approximation is that structure prediction models answer: if these chains occurred together in a Protein Data Bank (PDB) entry, what would their complex look like? Generative models answer a related question: given a target chain in a complex, what kinds of chains would be in contact with it? Both questions assume the chains appear together in a PDB structure. As a result, model confidence metrics such as ipTM estimate structural error assuming the true structure is in PDB; they are not calibrated, general-purpose measurements of interaction probability or affinity.

Distribution of AF3 ipTM scores for experimentally validated binders and non-binders across designed proteins targeting 15 targets.
Experimentally validated binders and likely non-binders overlap substantially in AF3 ipTM (confirmed binders in red, likely non binders in grey). Structural confidence can enrich for successful designs in some settings, but does not cleanly separate the classes. Data and public code from [5].

Antibody–antigen docking accuracy from the IsoDDE technical report.
IsoDDE demonstrates major progress in docking known cognate antibody–antigen partners. This is a different task from determining whether arbitrary antibody–antigen pairs bind. [6]

Published evaluations illustrate several related limitations:

  • Native cofolding confidence can enrich for successful designs in some pipelines, but is not a generally calibrated predictor of affinity[5][7].
  • Models that accurately dock known cognate partners may still struggle to distinguish cognate from non-cognate pairs[8].
  • Predicted structures and confidence remain nearly invariant under extensive mutations, including mutations that may alter folding or function[9].
  • Standard model outputs are static structures, not validated predictions of kinetic rates, equilibrium populations, or transitions between physical states[10].

These limitations do not diminish the progress in structure prediction and design. In fact, it is impressive that structural models are already so useful for binder design. But they expose a mismatch between the supervision currently available and the properties many downstream applications require.

Static structural data has taken the field extraordinarily far. Calibrated prediction of binding, mutation effects, affinity, kinetics, and specificity will require a qualitatively different class of experimental data.

What data would close these gaps?

The largest obvious source of additional data for protein models comes from protein sequence databases like UniProt. Sequence pretraining can substantially improve protein representations and additionally encodes indirect information about function, stability, and interaction, as demonstrated by ESMFold2[11]. But sequence co-evolution does not directly provide standardized labels for kon, koff, affinity, or therapeutic cross-reactivity. Indeed, the current best structure prediction model for antibody-antigen docking (OpenDDE) doesn't use protein language models[12].

The ideal dataset for training models to predict protein interaction kinetics would instead contain direct measurements of protein-protein interaction (PPI) kinetics. We think that requires scale, many binders measured against many targets, hard negatives, dense local mutation neighborhoods, direct kinetic labels, low noise, and a short enough experimental cycle for active learning.

Throughput and latency

The amount of data required for broad cross-target generalization is not known. These models probably have some latent understanding of protein biophysics that could be elicited by a small amount of PPI data – but it's very unlikely this would be enough to achieve the extreme levels of performance required. Our working hypothesis is that getting there will require at least hundreds of thousands of standardized measurements spanning many targets, binder scaffolds, mutations, and true negatives. More important than raw count may be the structure of the dataset: diverse target families, square cross-reactivity matrices, dense local mutation neighborhoods, and carefully controlled assay conditions.

Similarly, efficient model improvement and active learning require rapid iteration times: an ideal assay would have a latency measured in days or weeks.

Library on library

Existing benchmarks suggest that current models cannot reliably identify off-target interactions in advance[8]. Targeted screening based on model predictions may therefore miss the interactions that would provide the most useful corrective signal. Measuring the full interaction matrices between large collections of binders and targets avoids that selection problem. Screens containing many binders but only a few targets provide valuable information within those systems, but may offer limited evidence about cross-target generalization. Broader library-on-library measurements could complement them by exposing sparse and unexpected interactions.

Similarly, given that current performance varies widely across binders and targets, we'd need to see a large number of both targets and binders: models trained on many binders against a small number of targets won't necessarily generalize.

Kinetics rather than affinity

Ideally we'd measure the kinetics of each protein-protein interaction, rather than just the affinity. Very briefly, the kinetics parameters, kon and koff, are the rates at which two proteins form a complex and dissociate, respectively. Affinity (KD) is the ratio of the two, KD = koff / kon, and is typically used to summarize the strength of protein-protein interactions. Many scaled assays only really measure enrichment (typically a proxy for affinity), whereas lower-throughput techniques like BLI/SPR measure both kinetics parameters. Kinetics themselves are often therapeutically relevant[4], but more importantly they provide a more informative training target: kon and koff are influenced by different biophysical parameters. Additionally, protein-protein interactions are fundamentally about protein dynamics and structural ensembles; current structure predictors do not generally produce validated physical ensembles or transition rates[10].

Low-noise mutational data

Broad cross-target screens should be complemented by dense local perturbation data. Traditional alanine scans provide high-quality mutation effects but are limited in scale, whereas deep mutational scans provide broad sequence coverage but report enrichment rather than direct kinetic constants. The ideal training set would contain standardized measurements for dense single-mutant neighborhoods on many binders and targets, together with selected double-mutant panels that expose epistasis. This flavor of data would increase model sensitivity to local mutations and improve in silico maturation and search.

Is it possible to generate this data today?

Existing kinetics measurement techniques like BLI or SPR often cost tens to hundreds of dollars per data point, even at larger scales: generating billion-scale datasets is out of reach. This is still true for (very exciting!) emerging biosensor platforms such as SPOC that can collect thousands of kinetic measurements in parallel[13]. Pooled assays such as AlphaSeq and MP3-seq demonstrate a different and highly complementary advantage: large-scale interaction measurements at much lower cost[14][15]. That scale comes with tradeoffs. Their readouts can combine intrinsic binding with expression, valency, avidity, growth, and selection conditions, making it difficult to isolate the underlying biophysical interaction and potentially producing false negatives when, for example, one partner is poorly expressed. In particular, these assays generally do not provide independently calibrated, monovalent kon and koff measurements for every pair – the quantities we want as model-training targets.

These assays are quite valuable, especially for improving or validating designs within defined experimental systems. However, the published demonstrations above do not yet establish general prediction of kinetics and specificity across diverse protein families.

At Escalante, we had an idea: is it possible to collect kinetics data using DNA sequencing as a readout without compromising on quality? The answer, it turns out, is yes.

Escalante's experimental approach closes a key gap

These are the results that convinced us. The experiment we'll show was small in scale, only measured koff, and had dropouts. But it demonstrates that we can measure protein-protein interaction kinetics from sequencing reads alone – all in a single one-pot overnight reaction! We want to give the assay and the data the room they deserve in their own posts, but for now we want to share this first milestone.

For this campaign we generated 15 designs with mosaic[16] against a single target. Because our assay can handle multiple targets, we included 4 additional completely unrelated targets as likely negative controls, bringing the total number of protein-protein interactions measured in this experiment to 75. We tested them using our in-house assay and compared to third-party BLI measurements.

📝
Quick note: here we report kinetics as dwell time (1 / koff), the average time two proteins stay bound, since we think this is more intuitive than the clunkier biochemical koff, whose units are 1 / s. We'll also refer to q, a simple statistic we compute directly from our sequencing data that closely estimates koff.

A critical property of any assay are precision and accuracy: how close are your experimental replicates and how closely do your measurements match an accepted value (in our case BLI measurements of the same interactions), respectively? The figures display the corresponding estimated dwell time in seconds.

Estimated dwell times across two technical replicates, closely following the identity line.
(A) On-target estimated dwell times across two technical replicates. The estimates closely follow the identity line (Pearson r = 0.999; median fold difference = 1.04×).
Raw assay signal versus BLI-derived dwell time, with RANSAC calibration line.
(B) The overall RANSAC calibration [17] accurately recovers on-target dissociation rates for 10 of 13 binders with quantitative BLI references (LOOCV Pearson r = 0.972; median fold error = 1.15×).

Thirteen of the 15 designs had quantitative BLI koff fits. Our assay detected 10 of those 13 interactions, corresponding to a dropout rate of 23%. Among the 10 detected pairs, a leave-one-out calibration gave Pearson r = 0.972 and a median absolute fold error of 1.15 for koff compared to BLI. While it's too early to establish assay-wide properties, these results suggest we're much closer in accuracy to biophysical methods like BLI[[4]] and have the potential to match high-throughput methods in scale.

Our assay additionally estimates dwell times for the (presumed negative) off-targets. Encouragingly these are largely much shorter than the on-target dwell times.

15 by 5 matrix of estimated dwell times for each binder against the on-target and four control targets.
15 × 5 estimated dwell-time matrix. The on-target interaction has the longest estimated dwell time for 12 of 15 binders, with a median 4.8× separation from the longest-lived estimated off-target interaction.

What's next?

In this post we shared results from the experiment that convinced us our assay could deliver accurate and precise kinetic measurements of protein interactions. We demonstrated that you can reduce the typical O(n2) workflows for kinetics measurements down to O(1) through thoughtful assay design and the utility of DNA sequencing[18].

Over the next few posts we'll expand on the capabilities of the assay, how it works, and the role we think it will play in protein design as a fundamental measurement technology. Protein-protein interactions are just the start. Reach out if you're interested in building with us: hello@escalante.bio.


[[1]]: The most important of these constraints is whether a designed binder has the intended effect in a real biological system. In a future post we'll discuss our plan to measure such effects.

[[2]]: Binding to the target without any deleterious binding to off-targets.

[[3]]: This is deliberately napkin math.

[[4]]: For reference, many single amino acid mutations change kinetics far more than 1.15×.

References

  1. Xaira Therapeutics. Progressable Binders: The Binders That Matter. 2025.
  2. Andreessen Horowitz. Drug Discovery Has No Magic Wands. 2025.
  3. Cameron, B. J., Gerry, A. B., Dukes, J., Harper, J. V., Kannan, V., Bianchi, F. C., et al. Identification of a Titin-Derived HLA-A1-Presented Peptide as a Cross-Reactive Target for Engineered MAGE A3-Directed T Cells. Science Translational Medicine 5(197), 197ra103 (2013). doi:10.1126/scitranslmed.3006034
  4. Chen, P., Bordeau, B. M., Zhang, W., Balthasar, J. P. Investigations of Influence of Antibody Binding Kinetics on Tumor Distribution and Anti-Tumor Efficacy. The AAPS Journal 27(4), 91 (2025). doi:10.1208/s12248-025-01076-z
  5. Overath, M. D., Rygaard, A. S. H., Jacobsen, C. P., Brasas, V., Morell, O., Sormanni, P., Jenkins, T. P. Predicting Experimental Success in De Novo Binder Design: A Meta-Analysis of 3,766 Experimentally Characterised Binders. bioRxiv (2025). doi:10.1101/2025.08.14.670059. Code: de_novo_binder_scoring.
  6. Isomorphic Labs Team. Accurate Predictions of Novel Biomolecular Interactions with IsoDDE. Isomorphic Labs technical report (2026). doi:10.5281/zenodo.18606681
  7. Cotet, T.-S., Krawczuk, I., Stocco, F., Ferruz, N., Gitter, A., et al. Crowdsourced Protein Design: Lessons from the Adaptyv EGFR Binder Competition. bioRxiv (2025). doi:10.1101/2025.04.17.648362
  8. Smorodina, E., Ali, M., Kropivšek Brumat, K., Salicari, L., Miklavc, S., Kappassov, A., Fu, C., Sormanni, P., de Marco, A., Greiff, V. Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody–Antigen Structure Prediction Confidence Scores. bioRxiv (2026). doi:10.64898/2026.03.02.709004
  9. Feldman, J., Brogi, M., Skolnick, J. Adversarial Sequence Mutations in AlphaFold and ESMFold Reveal Nonphysical Structural Invariance, Confidence Failures, and Concerns for Protein Design. bioRxiv (2026). doi:10.64898/2026.02.25.708002
  10. Chakravarty, D., Schafer, J. W., Chen, E. A., Thole, J. F., Ronish, L. A., Lee, M., Porter, L. L. AlphaFold Predictions of Fold-Switched Conformations Are Driven by Structure Memorization. Nature Communications 15, 7296 (2024). doi:10.1038/s41467-024-51801-z
  11. Candido, S., Hayes, T., Derry, A., Rao, R., Lin, Z., Verkuil, R., et al. Language Modeling Materializes a World Model of Protein Biology. bioRxiv (2026). doi:10.64898/2026.06.03.729735
  12. CoFold Arena. CoFold Arena Leaderboard. Accessed 2026-08-11.
  13. Agu, C. V., Cook, R. L., Martelly, W., Gushgari, L. R., Mohan, M., et al. Multiplexed Proteomic Biosensor Platform for Label-Free Real-Time Simultaneous Kinetic Screening of Thousands of Protein Interactions. Communications Biology 8, 468 (2025). doi:10.1038/s42003-025-07844-z
  14. Younger, D., Berger, S., Baker, D., Klavins, E. High-Throughput Characterization of Protein–Protein Interactions by Reprogramming Yeast Mating. PNAS 114(46), 12166–12171 (2017). doi:10.1073/pnas.1705867114
  15. Baryshev, A., La Fleur, A., Groves, B., Michel, C., Baker, D., Ljubetič, A., Seelig, G. Massively Parallel Measurement of Protein–Protein Interactions by Sequencing Using MP3-seq. Nature Chemical Biology 20, 1514–1523 (2024). doi:10.1038/s41589-024-01718-x
  16. Escalante. mosaic: Composite-Objective Protein Design.
  17. Fischler, M. A., Bolles, R. C. Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography. Communications of the ACM 24(6), 381–395 (1981). doi:10.1145/358669.358692
  18. Escalante. Your Experiment Has a Runtime.