We docked every single product in a combinatorial reaction space (all 10,704
of them) to find out whether fragment-based synthon screening actually finds
the top scorers while docking only a fraction of the library. Short version:
the funnel machinery works, the published selection recipe didn’t survive the
scorer swap, and a one-line fix found 53 of the true top 100 compounds from
7.7% of the docking.
Make-on-demand combinatorial libraries are absurdly large. The Enamine REAL
space behind the V-SYNTHES2 method (Nazarova et al., *npj Drug Discovery*
2026) holds tens of billions of compounds. Nobody docks that. V-SYNTHES2’s
trick is a funnel: instead of docking products, dock a Minimal Enumeration
Library (MEL), which is every building-block fragment with its attachment
points capped by tiny placeholder groups. Then, geometrically, pick the
capped fragments that sit in the pocket with room to grow (”CapSelect”),
and only enumerate and dock the products of those winners. The paper reports
roughly a 10,000x compute reduction with enrichment factors in the hundreds.
We rebuilt the whole stack on Pauling: MEL generation, CapSelect, synthon
enumeration, and post-processing, all as Beam pipelines on a GPU Flink
cluster, with UniDock as the engine and full fragment-to-product lineage in
Postgres. Then we did something the paper’s scale doesn’t allow: we ran it on
a reaction space small enough to dock every product, so we could score the
funnel’s picks against perfect ground truth.
The experiment
Here’s the setup:
Target: Taken from PDB, with a P2Rank-selected pocket, run through our standard agent-driven receptor prep chain.
Space: one REAL reaction (`r_11`, an amide coupling) with about 100 synthons per position. The MEL is 209 capped fragments. Full enume
tiongives 10,900 unique products.Protocol: every molecule goes through the same deterministic prep (RDKit ETKDGv3 conformers, protonation at pH 7.4, MMFF94 minimization, Meeko PDBQT), then UniDock. The full-space run posed 10,704 of 10,900 compounds (98.2%).
Ground truth: the complete docking-scored space. So we can count, exactly, how many of the true top-100 and top-500 compounds any selection strategy catches.
One substitution up front: Vina, not ICM
The paper docked with ICM, whose sampler energy-minimizes every pose in a
full-atom force field. We run UniDock, and GNINA as well, and both rank with
Vina-style scoring on precomputed grids. We did test the obvious upgrade
before settling: re-docking our rehearsal set with GNINA’s CNN rescoring gave
worse pose recovery (all-pairs median 5.97 Å vs UniDock’s 4.44 Å) at twice
the cost, with the same 84% headline pose consistency. So Vina ranking it is,
and everything below should be read through that lens.
Caveat #1: the paper’s recipe did worse than random on our scorer
We first ran CapSelect exactly as published: same sphere-growth geometry,
same MergedScore weights balancing docking score against growth capacity. It
picked 19 fragments, and enumerating those gives 1,983 products, or 18.5% of
the space. Against ground truth:
| | Docked share | Caught of true top-100 | Caught of true top-500 | EF100 |
| ---------------------------- | ----------------------- | ---------------------- | ---------------------- | -------- |
| Full space (ground truth) | 100% (10,704 of 10,704) | 100 | 500 | - |
| Random (expectation) | 18.5% (1,983 of 10,704) | 18 | 92 | 1.0 |
| **CapSelect, paper weights** | 18.5% (1,983 of 10,704) | **5** | **33** | **0.65** |That’s not “no enrichment.” That’s anti-enrichment: dock 18.5% of the space
via the paper’s rule and you end up with fewer of the best compounds than
if you’d picked that 18.5% at random.
The mechanism is intuitive. The MergedScore weights deliberately
spend docking score to buy growth capacity (the selected fragments docked at
-5.63 median vs the pool’s -6.18), and their products come out smaller and
less greasy than the space at large (median MW 355 vs 400, logP 2.9 vs 3.6).
Vina-family scoring rewards size and grease. So on our scorer, the
growth-capacity trade reads as a score handicap. Those weights were tuned on
ICM scores; transplanted onto Vina ranking, they select against what the
scorer calls best.
The premise holds, and a one-line fix delivers the headline
The core assumption of the method is that a fragment’s docking score predicts
its products’ scores. Across all 200 docked parent fragments in the space:
One dot per MEL fragment: its UniDock docking score vs the median score of
its enumerated products (full r₁₁ space: 200 fragments, 10,704 docked
products). The paper-weight picks and the score-ranked picks landed on very
different parts of the band.
| Fragment score vs. its products' scores | Spearman ρ |
| --------------------------------------- | ---------- |
| Products' **median** score | **0.91** |
| Products' **best** score | 0.57 |That is a near-perfect ranking signal. So we ran the counterfactual: same
funnel, same 19-fragment budget, but ranked by fragment docking score alone,
with no growth-capacity term. Enumerating those 19 fragments gives 825
products. That’s 7.7% of the space, one in thirteen compounds:
| | Docked share | Caught of true top-100 | Caught of true top-500 | EF100 |
| ------------------------- | ------------------------ | ---------------------- | ---------------------- | ------- |
| Full space (ground truth) | 100% (10,704 of 10,704) | 100 | 500 | - |
| Random (expectation) | 7.7% (825 of 10,704) | 8 | 38 | 1.0 |
| **Score-ranked funnel** | **7.7% (825 of 10,704)** | **53** | **190** | **6.1** |
Half the top-100 compounds, found by docking 7.7% of the space. That’s
6.7x what random selection gets at the same budget, and for the record, this
selection overlaps the paper-weight selection by exactly one fragment. Under
the conservative property-matched control (random draws stratified to the
same molecular-weight and logP profile, which strips out the
size/lipophilicity edge a score-based rule inherits), enrichment is still
1.69x at top-100 and 1.76x at top-500. If you also count the 209 fragment
docks, the funnel’s total budget is about 10% of full enumeration, so roughly
a 10x reduction even at toy scale.
Of the K best-scoring compounds in the fully-docked space, how many each
selection pool contains. Solid lines are the funnels; dashed lines are what
a random pool of the same size expects. At K = 100, the score-ranked pool
holds 53 of the 100 best; the paper-weight pool holds 5.
The supporting machinery held up too. Every fragment’s binding mode is
recovered by at least one of its products (84% of fragments have a
pose-preserving product, and the best product per fragment sits at a median
RMSD of 0.21 Å), even though individual products scatter quite a bit
(all-pairs median 4.44 Å, which we already knew is a Vina-family pose
property per the GNINA test above).
Caveat #2: scale. Would 10x or 100x more data make this better?
Honest answer: we don’t know yet, and we’d rather leave it as an open
question than wave our hands. The mechanics cut both ways.
Why EF should grow: enrichment is the gap between what a signal-driven
selection catches and what random catches at the same budget. Random’s
expected catch of the top-K shrinks linearly as the space grows relative to a
fixed budget. At our scale, a random 7.7% draw still finds about 8 of the top
100. At REAL scale, a comparable budget covers about 0.003% of the space, and
random finds essentially nothing. With the fragment-to-product correlation
sitting at ρ = 0.91, the selected pool’s catch scales with the signal, not
with the coverage fraction. That’s the regime where the paper’s EF100 numbers
reach into the hundreds.
Why it might not be so simple: with thousands of synthons per position
(the paper’s regime), each fragment’s product cloud gets bigger, and the
best-product correlation (already diluted to ρ = 0.57 here) starts to matter
more. Products get bigger too, where our Vina scorer’s size/grease bias,
exactly the thing the property-matched control corrects for, becomes a bigger
confound in what “top-scoring” even means. And the growth-capacity term,
which anti-selected here, is designed for that regime. Give it thousands of
synthons to spend growth on and the trade may start paying, which would make
the right answer some recalibrated blend rather than pure score ranking.
So the next experiment is unambiguous: a licensed REAL reaction sample at 10x
to 100x the synthon count, with MergedScore weights recalibrated for
Vina-scale scores on this same benchmark. The platform is ready.



