Closed loop

Learning from the misses: post-assay error analysis in a closed design loop

Negative results carry most of the information in a design cycle, and almost all of it is lost if a flat sensorgram is recorded as a non-binder. The work is in deciding which kind of failure you are looking at.

A design cycle ends with a plate of candidates and a set of binding measurements. The successes are pleasant and relatively uninformative: they confirm that part of the model was right. The failures are where the next cycle comes from, and they are routinely wasted, because a candidate that produced no measurable signal gets filed under a single label when it actually belongs to one of at least four distinct categories.

Separating those categories is the step we spend the most care on. It is also the step that determines whether a loop is genuinely closed or merely iterative, in the sense that the same mistake is made again in a new sequence.

Four ways to produce a flat trace

The first is the one everyone assumes: the molecule does not bind the target. The second is that it binds but the measurement could not see it, because of an assay artifact. The third is that the sequence is correct but the molecule was not in its active conformation when it reached the instrument. The fourth is that the material was not what the label said, through synthesis failure, truncation, or degradation in storage.

These four demand completely different responses. The first is a real update to the structure-activity model. The second is an instrument problem and should not touch the model at all. The third is a protocol problem that will recur across every candidate in the batch. The fourth is a materials problem. Conflating them injects false negatives into training data, which is more damaging than having no data, because the model learns a confident wrong boundary.

A false negative recorded as a true negative does not just waste one experiment. It teaches the next generation of designs to avoid a region of sequence space that was never actually tested.

Artifacts that look like results

Surface plasmon resonance and biolayer interferometry are the workhorses of this stage, and both have characteristic failure modes that are well documented and still regularly misread.

Mass transport limitation appears when ligand density is high and analyte diffusion to the surface becomes rate-limiting. The fitted association rate is depressed, the dissociation phase is distorted by rebinding, and the result is a kinetic profile that is a property of the surface rather than of the interaction. The diagnostic is straightforward: vary flow rate and ligand density, and see whether the fitted constants move.

Avidity arises when the immobilized partner presents multiple binding sites, or when immobilization chemistry creates clusters. Apparent affinity improves, sometimes dramatically, and the improvement does not transfer to a solution-phase measurement. Orientation matters too: random amine coupling will inactivate a fraction of the target and may bury the epitope of interest entirely, which produces exactly the flat trace that gets filed as a non-binder.

Nonspecific and surface binding is the mirror image, producing signal where there is no specific interaction. Phosphorothioate-containing backbones are particularly prone to it. Reference-channel subtraction catches some of this; a scrambled-sequence control of matched length and chemistry catches more.

Aggregation gives superstoichiometric responses, poor fits, and non-reproducible kinetics. Regeneration damage degrades the surface across a run, so late-plate candidates systematically underperform early ones. Buffer mismatch between the sample and the running buffer produces bulk refractive index shifts that can swamp a weak interaction. Each of these has a plate-level signature, which is why we analyze batches rather than individual wells.

Folding is a protocol variable

An aptamer that has not been refolded is a different molecule from the one that was designed. Our standard requires a defined annealing step and a recorded free magnesium concentration for every measurement, because divalent cation dependence is strong enough that an otherwise identical candidate can look inactive at plasma-relevant magnesium and excellent at 5 millimolar.

When an entire batch underperforms its predictions, folding protocol is the first hypothesis, not the last. The signature is a batch-wide shift rather than a candidate-specific one, and it is distinguishable from a model failure precisely because it does not correlate with any sequence feature.

The triage table we actually use

ObservationLeading hypothesisNext experiment
Flat trace, control candidate also flatTarget inactive or epitope occluded by immobilizationRe-immobilize with oriented capture; confirm with a known binder
Flat trace, controls behave normallyTrue non-binderRecord as a negative with full confidence weight
Whole batch weaker than predictedRefolding or magnesium protocolRefold series across a magnesium range on a subset
Fitted on-rate falls as ligand density risesMass transport limitationReduce ligand density; vary flow rate
Affinity much better on surface than in solutionAviditySolution-phase competition measurement
Signal with scrambled control of matched chemistryNonspecific surface bindingBlocking optimization; reconsider backbone chemistry
Poor fit, superstoichiometric responseAggregationSize-exclusion or light-scattering check on the sample
Late wells systematically weakerSurface degradation from regenerationMilder regeneration; randomize plate order
Mass spectrometry inconsistent with designSynthesis or storage failureResynthesize; exclude from model update entirely

None of this is novel. All of it is well known to anyone who runs these instruments for a living. The difference a design loop makes is that the triage happens systematically on every batch, the conclusion is recorded in a structured form, and the resulting confidence weight travels with the data point into the next model update.

What the model update actually looks like

Once failures are classified, the cycle produces three separate updates, and keeping them separate is what prevents the loop from thrashing.

The first is to the structure-activity hypothesis. Confirmed non-binders that were predicted to bind are the valuable cases. We look for shared features: a particular loop length, a stem that the ensemble analysis says is marginally stable, a modification at a position that turns out to be interface-proximal. Where a cluster emerges, the constraint specification changes, and the change is written down with its rationale so that a later cycle can test whether it was correct.

The second is to model calibration. We stratify outcomes by the confidence the model assigned beforehand. A model whose high-confidence predictions fail at the same rate as its low-confidence ones is not providing usable ranking information, regardless of its aggregate accuracy. That diagnosis changes the acquisition strategy rather than the chemistry.

The third is to batch selection. The next set of candidates is chosen for expected information gain, not for best predicted affinity, which in practice means deliberately including designs that sit near a decision boundary the loop needs resolved. This is ordinary active learning, and it is the reason a cycle can be informative even when its hit rate is low.

cycle n · triage summary G-rich stem variants predicted bind · flat → artifact (occluded epitope, re-run) short-loop series predicted bind · flat → confirmed non-binder, constraint updated C-rich interface variants predicted bind · bound → retained, truncation series queued batch-wide shift all candidates weak → refolding protocol, excluded from update

Our wet-lab work runs in collaboration with laboratories at KAIST, which is what makes this cadence possible: the triage conversation happens with the people who ran the instrument, while the plate is still fresh and the raw sensorgrams are still at hand. Reasoning agents prepare the analysis, flag the artifact signatures, and draft the proposed constraint changes; the scientists who generated the data decide which of those drafts is right. Post-assay error analysis is not a step that automates cleanly, and it is not a step that scales without help either.

Further reading

  • Myszka, Current Opinion in Biotechnology 1997, on kinetic analysis with surface plasmon resonance biosensors, including mass transport effects.
  • Rich and Myszka, in their long series of annual reviews of commercial biosensor literature, on recurring experimental design errors.
  • Keefe, Pai and Ellington, Nature Reviews Drug Discovery 2010, for aptamer characterization practice.
  • Lorenz and colleagues, Algorithms for Molecular Biology 2011, ViennaRNA 2.0, for the ensemble calculations referenced in folding triage.
  • Chen and colleagues, 2022, on RNA-FM and foundation-model representations of RNA sequence, for the modeling side of the loop.
  • Hoinka and colleagues, on AptaSUITE, for treating high-throughput selection sequencing as noisy supervision.

← All insights