KinaseFoundationModel Version 2 Limitations Download the models Example campaign · unlisted

An example campaign: three models, each doing a job the others cannot

One worked pass over a real catalogue, run in selection rather than in the laboratory. It is the only artefact that puts all three model families in a single workflow. The Methods section below, from the original selection report, gives each step in the order it happened; what follows here is the context that report assumes.

Why three models and not one. The KKB per-target random forests can score millions of compounds cheaply, but they answer is this active here one target at a time, so they cannot rank two compounds reliably against each other and cannot compare two kinases at all. The two version 2 comparators answer exactly those two questions, but they are far too slow to run over five million compounds. Putting the cheap filter first and the comparators second is what makes the campaign tractable: 86 million predictions narrow the field, then 20 slots are chosen by comparison.

Stage one, the shortlist. Every KKB release ships two models per kinase, built from the curated SAR in that release: 392 activity classifiers at a median ROC-AUC of 0.914, rising to 0.957 across the 95 best-studied targets, and 319 potency regressors at a median error of 0.48 log units, with 88% of predictions inside one log. Those models scored all 4,774,663 Enamine compounds against all nine kinases and handed over the top 100,000 per kinase. They shortlist and nothing more. Full performance of the KKB per-target models →

Stages two and three, the choosing. LigASeqLigB places each survivor against a measured reference ladder and scores it by the fraction of that ladder it beats; SeqALigSeqB then decides, for each paralog pair, which sibling the compound prefers. The detail of both, including the reference ladders and why pan-kinase compounds are barred, is in the Methods below. How the two comparators work →

What a run would settle. Every model in the chain was fitted on training data that is 65 to 97 per cent active, so none has seen enough inactive chemistry to predict inactivity, and 88% of the predictions here sit above the activity threshold for that reason alone. The survivors also sit at a median Tanimoto of 0.27 to 0.36 from the chemistry each model was fitted on, which is the distant-and-related regime where measured accuracy is weakest. Both facts point the same way: a real screen supplies the negatives and the novel chemistry the models lack, which is why this panel is worth running rather than merely worth reading.


Kinase screening panel

Enamine screening collection, 4,774,663 compounds · 9 understudied kinase targets · selection completed 11 August 2026
Selection complete. 20 compounds across 9 targets.
LIMK1
2
pIC50 6.42 to 6.51
LIMK2
2
pIC50 7.67 to 7.86
CLK1
4
pIC50 7.02 to 7.61
CLK2
2
pIC50 6.64 to 6.69
CLK4
4
pIC50 7.30 to 8.06
MAP4K4
1
pIC50 8.57 to 8.57
TAOK1
2
pIC50 6.26 to 6.43
TAOK3
2
pIC50 7.11 to 7.25
GAK
1
pIC50 8.03 to 8.03

Methods

How these 20 compounds were chosen, in the order it happened.
1
Screen everything, then hand over
All 4,774,663 compounds in the Enamine screening collection scored against all 9 kinases by the KKB per-target random forests, regressor and classifier: 86 million predictions, nothing sampled. These models only shortlist, narrowing to the top 100,000 candidates per kinase. They do not choose the panel.
2
v2 potency: position on a measured ladder
For each kinase, 10 to 16 real measured compounds form a reference ladder, starting near pIC50 5.0 and topping out between 7.3 and 9.0 depending on how much measured chemistry the kinase has. LigASeqLigB v2 [ligand | sequence | ligand] compares every candidate against every rung; its score is the fraction of the ladder it beats. Validation: it ordered 24 of 28 known LIMK2 pairs correctly (86%).
2b
Pan-kinase compounds are excluded
Any compound measured at pIC50 ≥ 7 on more than 20 kinases in KKB is barred, both as a ladder rung and from the panel itself, 183 skeletons in all. Staurosporine alone is measured potent on 334 kinases; left in, it makes “beats known chemistry” mean “beats a compound that beats everything”, and it cannot support a selectivity claim.
3
v2 selectivity: which sibling does it prefer
SeqALigSeqB v2 [sequence | ligand | sequence] is asked, for each paralog pair, which kinase the compound prefers. Only positive calls are kept, both directions of a pair are filled wherever the model will commit, and chemical dissimilarity is enforced so each slot is a distinct scaffold. The panel is built to tell sibling kinases apart, not to maximise raw potency. 4 compounds are marked exploratory: directions where the model reaches no confident call, kept deliberately because that is where measurement would teach us most.
Predicted pIC50: potency on the standard log scale, where 6.0 is 1 µM, 7.0 is 100 nM, 8.0 is 10 nM.
v2 potency: the fraction of that kinase's measured reference ladder the compound is predicted to beat. 1.000 means it is placed above every known compound on the ladder, 0.500 means mid-scale.
v2 prefers: the probability the selectivity model assigns to the compound preferring its primary kinase over the named sibling. Above 0.5 is a positive call; 0.95 is a strong one. This is what a single assay can confirm or refute.
Exploratory: a slot where the selectivity model reaches no confident call in that direction. Kept on purpose: those are the sibling pairs with almost no measured data, so a real number would be worth the most.
Upper heatmap: each compound minus its own mean across the 9 kinases. Red where it is relatively strong, blue where relatively weak. The absolute scale cannot show this: 88% of these predictions sit above the activity threshold, because these models are trained on cohorts that are 65 to 97 percent active and cannot express inactivity.
Lower heatmap: the same matrix on the raw pIC50 scale.
Purchase list · 20 compounds, Enamine catalogue IDs, SMILES, predicted potency on all 9 targets, the v2 ladder win rate and the v2 sibling preference for every discrimination slot.
Download CSV (20 compounds)

Selectivity profile · v2 model selection

Each compound relative to its own mean across the 9 targets. Red = the targets that compound prefers, blue = the targets it avoids. This is the profile a screen would produce, and it is the view that shows selectivity.
−1.2 → 0 → +1.2 log units vs the compound's own mean
LIMK1
LIMK2
CLK1
CLK2
CLK4
MAP4K4
TAOK1
TAOK3
GAK
1 CLK4 Z3337539906
NC(c1nc(-c2c[nH]c3ncccc23)cs1)C1CC1
-0.80-0.06+0.39+0.40+1.18+0.72-0.94-0.61-0.28
2 LIMK2 Z68319703
CCN(Cc1ccccc1)C(=O)c1ccc(S(=O)(=O)Nc2ccccc2)cc1
+0.16+1.31+0.31-0.34-0.04-0.61-0.88+0.27-0.19
3 TAOK3 Z9898331572
CCCCCCOc1ccc(C(=O)NCC(=O)NC2CCCc3ccccc32)cc1
-0.46-0.15+0.20-0.47-0.13-0.37+0.66+0.74-0.06
4 TAOK1 Z2216894329
CC(=O)c1c(C)c2cnc(Nc3ccc(N4CCNCC4)cn3)nc2n(C2CCCC2)c1=O
+0.20-0.20+0.14+0.21+0.91-0.31-0.13-0.82+0.04
5 CLK4 Z46624327
COc1ccc(C=C2SC(=NC3CCCC3)NC2=O)cc1O
-0.74+0.03+0.36+0.09+0.93-0.19-0.84+0.14+0.23
6 LIMK1 Z838463180
Cc1ncc2c(n1)CCC(NC(=O)N1CC=C(c3c(C)[nH]c4ccccc34)CC1)C2
+0.13+0.13+0.23-0.21+0.07-0.12-0.55+0.16+0.16
7 CLK1 Z4082010706
Cc1n[nH]c2ccc(-c3cnc(N(C)C4CCNCC4)cn3)cc12
-0.44-0.18+0.81+0.12+0.83+0.05-0.84-0.41+0.04
8 CLK2 Z7684579660
Cc1nnc(CNc2nc(N[C@@H]3CCCO[C@H]3c3ccc(Cl)cc3)c3ncn(C(C)C)c3n2)n1C1CC1
-0.06+0.01+0.23+0.26+0.19-0.06-0.45-0.31+0.23
9 CLK1 Z5129795909
C[C@@H](c1ccccc1Br)n1cc(-c2cnn3c2CN(C(=O)c2cc(-c4ccc5c(c4)OCCO5)n[nH]2)CC3)nn1
-0.09+0.22+0.61+0.15+0.47-0.21-0.61-0.35-0.15
10 MAP4K4 Z2788059508
CC(C)(Oc1ccc(-c2cnc(N)c(-c3ccc(Cl)cc3)c2)cc1)C(=O)O
-0.60+0.14+0.16-0.48-0.16+2.02-0.82-0.20-0.04
11 GAK Z2568724659
CCOc1cc2ncc(C#N)c(Nc3ccc(F)c(Cl)c3)c2cc1NC(=O)/C=C/CN(C)C
-0.39+0.40-0.29-0.29-0.56+0.11-0.21-0.37+1.63
12 CLK4 Z4082008670
c1cnc2[nH]cc(-c3ccnc(NC4CCNC4)n3)c2c1
-0.73-0.30+0.32+0.70+1.28+0.66-0.74-0.77-0.46
13 LIMK2 Z8999112634
CC(C)C(=O)Nc1ncc(C(=O)NCCN(Cc2ccccc2)C(=O)c2ccc(S(=O)(=O)Nc3ccccc3)cc2)s1
+1.29+0.94-0.14-0.44-0.03-0.55-0.50-0.30-0.31
14 TAOK3 Z2568721748
Cc1[nH]c(/C=C2\C(=O)Nc3ccc(S(=O)(=O)Cc4c(Cl)cccc4Cl)cc32)c(C)c1C(=O)N1CCC[C@@H]1CN1CCCC1
-0.36-0.03-0.24-0.01+0.07+0.18-0.14+0.67-0.11
15 TAOK1 Z2216912400
CN(c1cccc(CNc2nc(Nc3ccc4c(c3)CC(=O)N4)ncc2C(F)(F)F)c1)S(C)(=O)=O
-0.07+0.13+0.18+0.04+0.46+0.06+0.08-0.86+0.02
16 CLK4 Z44301537
CCOc1cc(C=C2SC(=S)NC2=O)ccc1O
-0.72+0.25+0.16+0.17+1.08-0.40-0.80+0.19+0.03
17 LIMK1 Z2612280147
Cc1[nH]c2ccccc2c1C1=CCN(c2cnc(N3CCCC3=O)cn2)CC1
+0.12+0.19+0.26+0.02+0.10-0.22-0.69-0.18+0.38
18 CLK1 Z89616206
COc1ccc2cc(CN(C)C(=O)c3cc4cc([N+](=O)[O-])ccc4s3)ccc2c1
-0.63+0.05+1.05+0.21+0.66-0.54-0.79-0.01+0.00
19 CLK2 Z7684579671
CC(C)CCOCCNc1nc(NC2COC(C3CC3)C2)c2ncn(C(C)C)c2n1
-0.03+0.10+0.31+0.25+0.22-0.22-0.36-0.30+0.06
20 CLK1 Z1918130497
Cc1cc(-c2cc(-c3nc(-c4cnc(N(C)C)cn4)no3)c3cnn(C(C)C)c3n2)c(C)o1
-0.30+0.23+0.52+0.17+0.33-0.39-0.44-0.13-0.03

Absolute predicted potency

The same matrix on the raw pIC50 scale, centred at 7.0 rather than the 6.0 activity threshold. 88% of cells sit above 6.0, so an activity-threshold scale paints almost everything red: these regressors are trained on cohorts that are 65 to 97 percent active and cannot express inactivity. That limitation is the reason the panel is worth buying, since profiling supplies the negatives the models lack.
5.5 → 7.0 → 8.6 predicted pIC50
LIMK1
LIMK2
CLK1
CLK2
CLK4
MAP4K4
TAOK1
TAOK3
GAK
1 CLK4 Z3337539906
NC(c1nc(-c2c[nH]c3ncccc23)cs1)C1CC1
5.896.637.087.097.877.415.756.086.41
2 LIMK2 Z68319703
CCN(Cc1ccccc1)C(=O)c1ccc(S(=O)(=O)Nc2ccccc2)cc1
6.717.866.866.216.515.945.676.826.36
3 TAOK3 Z9898331572
CCCCCCOc1ccc(C(=O)NCC(=O)NC2CCCc3ccccc32)cc1
6.056.366.716.046.386.147.177.256.45
4 TAOK1 Z2216894329
CC(=O)c1c(C)c2cnc(Nc3ccc(N4CCNCC4)cn3)nc2n(C2CCCC2)c1=O
6.596.196.536.607.306.086.265.576.43
5 CLK4 Z46624327
COc1ccc(C=C2SC(=NC3CCCC3)NC2=O)cc1O
5.636.406.736.467.306.185.536.516.60
6 LIMK1 Z838463180
Cc1ncc2c(n1)CCC(NC(=O)N1CC=C(c3c(C)[nH]c4ccccc34)CC1)C2
6.516.516.616.176.456.265.836.546.54
7 CLK1 Z4082010706
Cc1n[nH]c2ccc(-c3cnc(N(C)C4CCNCC4)cn3)cc12
6.096.357.346.657.366.585.696.126.57
8 CLK2 Z7684579660
Cc1nnc(CNc2nc(N[C@@H]3CCCO[C@H]3c3ccc(Cl)cc3)c3ncn(C(C)C)c3n2)n1C1CC1
6.376.446.666.696.626.375.986.126.66
9 CLK1 Z5129795909
C[C@@H](c1ccccc1Br)n1cc(-c2cnn3c2CN(C(=O)c2cc(-c4ccc5c(c4)OCCO5)n[nH]2)CC3)nn1
6.416.727.116.656.976.295.896.156.35
10 MAP4K4 Z2788059508
CC(C)(Oc1ccc(-c2cnc(N)c(-c3ccc(Cl)cc3)c2)cc1)C(=O)O
5.956.696.716.076.398.575.736.356.51
11 GAK Z2568724659
CCOc1cc2ncc(C#N)c(Nc3ccc(F)c(Cl)c3)c2cc1NC(=O)/C=C/CN(C)C
6.016.806.116.115.846.516.196.038.03
12 CLK4 Z4082008670
c1cnc2[nH]cc(-c3ccnc(NC4CCNC4)n3)c2c1
6.056.487.107.488.067.446.046.016.32
13 LIMK2 Z8999112634
CC(C)C(=O)Nc1ncc(C(=O)NCCN(Cc2ccccc2)C(=O)c2ccc(S(=O)(=O)Nc3ccccc3)cc2)s1
8.027.676.596.296.706.186.236.436.42
14 TAOK3 Z2568721748
Cc1[nH]c(/C=C2\C(=O)Nc3ccc(S(=O)(=O)Cc4c(Cl)cccc4Cl)cc32)c(C)c1C(=O)N1CCC[C@@H]1CN1CCCC1
6.086.416.206.436.516.626.307.116.33
15 TAOK1 Z2216912400
CN(c1cccc(CNc2nc(Nc3ccc4c(c3)CC(=O)N4)ncc2C(F)(F)F)c1)S(C)(=O)=O
6.286.486.536.396.816.416.435.496.37
16 CLK4 Z44301537
CCOc1cc(C=C2SC(=S)NC2=O)ccc1O
5.606.576.486.497.405.925.526.516.35
17 LIMK1 Z2612280147
Cc1[nH]c2ccccc2c1C1=CCN(c2cnc(N3CCCC3=O)cn2)CC1
6.426.496.566.326.406.085.616.126.68
18 CLK1 Z89616206
COc1ccc2cc(CN(C)C(=O)c3cc4cc([N+](=O)[O-])ccc4s3)ccc2c1
5.936.617.616.777.226.025.776.556.56
19 CLK2 Z7684579671
CC(C)CCOCCNc1nc(NC2COC(C3CC3)C2)c2ncn(C(C)C)c2n1
6.366.496.706.646.616.176.036.096.45
20 CLK1 Z1918130497
Cc1cc(-c2cc(-c3nc(-c4cnc(N(C)C)cn4)no3)c3cnn(C(C)C)c3n2)c(C)o1
6.206.737.026.676.836.116.066.376.47

Chemical diversity of the panel

Ligand against ligand, 20 x 20, Morgan Tanimoto. Row and column numbers are the compound numbers used throughout this page. A diverse panel is mostly blue. Median off-diagonal similarity is 0.12, maximum 0.52, against a selection ceiling of 0.55. Red would mean two slots are near-duplicates and one is wasted.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
1 Z33375399060.060.100.130.110.150.190.140.140.090.070.460.100.090.140.090.140.060.120.11
2 Z683197030.060.210.100.100.110.060.110.100.120.150.060.520.150.150.110.110.170.080.07
3 Z98983315720.100.210.140.150.190.110.120.080.130.130.100.210.110.100.160.130.160.120.09
4 Z22168943290.130.100.140.150.190.180.180.110.110.140.180.110.160.160.130.180.080.150.13
5 Z466243270.110.100.150.150.110.120.070.100.100.100.080.090.170.110.440.100.140.080.05
6 Z8384631800.150.110.190.190.110.170.130.150.120.120.170.120.170.130.090.440.100.110.11
7 Z40820107060.190.060.110.180.120.170.110.170.130.130.200.090.120.160.100.170.080.090.22
8 Z76845796600.140.110.120.180.070.130.110.110.120.110.150.120.120.140.100.110.060.410.14
9 Z51297959090.140.100.080.110.100.150.170.110.120.120.130.130.150.130.090.130.090.100.13
10 Z27880595080.090.120.130.110.100.120.130.120.120.170.090.130.110.120.120.120.150.070.11
11 Z25687246590.070.150.130.140.100.120.130.110.120.170.080.160.120.170.170.080.160.070.14
12 Z40820086700.460.060.100.180.080.170.200.150.130.090.080.070.100.150.070.140.040.180.10
13 Z89991126340.100.520.210.110.090.120.090.120.130.130.160.070.130.170.090.130.170.110.10
14 Z25687217480.090.150.110.160.170.170.120.120.150.110.120.100.130.160.150.170.100.060.07
15 Z22169124000.140.150.100.160.110.130.160.140.130.120.170.150.170.160.120.150.110.120.10
16 Z443015370.090.110.160.130.440.090.100.100.090.120.170.070.090.150.120.110.120.100.09
17 Z26122801470.140.110.130.180.100.440.170.110.130.120.080.140.130.170.150.110.060.070.10
18 Z896162060.060.170.160.080.140.100.080.060.090.150.160.040.170.100.110.120.060.070.13
19 Z76845796710.120.080.120.150.080.110.090.410.100.070.070.180.110.060.120.100.070.070.11
20 Z19181304970.110.070.090.130.050.110.220.140.130.110.140.100.100.070.100.090.100.130.11

Panel compounds · v2 model selection · 20 compounds

CLK4CLK4 over CLK2
7.87predicted pIC50
v2 potency 1.0 of the reference ladder beaten · v2 prefers CLK4 over CLK2 at 0.937
Z3337539906
NC(c1nc(-c2c[nH]c3ncccc23)cs1)C1CC1
LIMK2LIMK2 over LIMK1
7.86predicted pIC50
v2 potency 0.938 of the reference ladder beaten · v2 prefers LIMK2 over LIMK1 at 0.885
Z68319703
CCN(Cc1ccccc1)C(=O)c1ccc(S(=O)(=O)Nc2ccccc2)cc1
TAOK3TAOK3 over TAOK1
7.25predicted pIC50
v2 potency 1.0 of the reference ladder beaten · v2 prefers TAOK3 over TAOK1 at 0.829
Z9898331572
CCCCCCOc1ccc(C(=O)NCC(=O)NC2CCCc3ccccc32)cc1
TAOK1TAOK1 over TAOK3
6.26predicted pIC50
v2 potency 0.571 of the reference ladder beaten · v2 prefers TAOK1 over TAOK3 at 0.731
Z2216894329
CC(=O)c1c(C)c2cnc(Nc3ccc(N4CCNCC4)cn3)nc2n(C2CCCC2)c1=O
CLK4CLK4 over CLK1
7.3predicted pIC50
v2 potency 1.0 of the reference ladder beaten · v2 prefers CLK4 over CLK1 at 0.724
Z46624327
COc1ccc(C=C2SC(=NC3CCCC3)NC2=O)cc1O
LIMK1LIMK1 over LIMK2
6.51predicted pIC50
v2 potency 1.0 of the reference ladder beaten · v2 prefers LIMK1 over LIMK2 at 0.720
Z838463180
Cc1ncc2c(n1)CCC(NC(=O)N1CC=C(c3c(C)[nH]c4ccccc34)CC1)C2
CLK1CLK1 over CLK2
7.34predicted pIC50
v2 potency 0.938 of the reference ladder beaten · v2 prefers CLK1 over CLK2 at 0.702
Z4082010706
Cc1n[nH]c2ccc(-c3cnc(N(C)C4CCNCC4)cn3)cc12
CLK2CLK2 over CLK1exploratory
6.69predicted pIC50
v2 potency 0.938 of the reference ladder beaten · v2 prefers CLK2 over CLK1 at 0.585
Z7684579660
Cc1nnc(CNc2nc(N[C@@H]3CCCO[C@H]3c3ccc(Cl)cc3)c3ncn(C(C)C)c3n2)n1C1CC1
CLK1CLK1 over CLK4exploratory
7.11predicted pIC50
v2 potency 0.938 of the reference ladder beaten · v2 prefers CLK1 over CLK4 at 0.563
Z5129795909
C[C@@H](c1ccccc1Br)n1cc(-c2cnn3c2CN(C(=O)c2cc(-c4ccc5c(c4)OCCO5)n[nH]2)CC3)nn1
MAP4K4MAP4K4 potency
8.57predicted pIC50
v2 potency 1.0 of the reference ladder beaten
Z2788059508
CC(C)(Oc1ccc(-c2cnc(N)c(-c3ccc(Cl)cc3)c2)cc1)C(=O)O
GAKGAK potency
8.03predicted pIC50
v2 potency 1.0 of the reference ladder beaten
Z2568724659
CCOc1cc2ncc(C#N)c(Nc3ccc(F)c(Cl)c3)c2cc1NC(=O)/C=C/CN(C)C
CLK4CLK4 over CLK2
8.06predicted pIC50
v2 potency 1.0 of the reference ladder beaten · v2 prefers CLK4 over CLK2 at 0.927
Z4082008670
c1cnc2[nH]cc(-c3ccnc(NC4CCNC4)n3)c2c1
LIMK2LIMK2 over LIMK1
7.67predicted pIC50
v2 potency 0.938 of the reference ladder beaten · v2 prefers LIMK2 over LIMK1 at 0.759
Z8999112634
CC(C)C(=O)Nc1ncc(C(=O)NCCN(Cc2ccccc2)C(=O)c2ccc(S(=O)(=O)Nc3ccccc3)cc2)s1
TAOK3TAOK3 over TAOK1
7.11predicted pIC50
v2 potency 1.0 of the reference ladder beaten · v2 prefers TAOK3 over TAOK1 at 0.780
Z2568721748
Cc1[nH]c(/C=C2\C(=O)Nc3ccc(S(=O)(=O)Cc4c(Cl)cccc4Cl)cc32)c(C)c1C(=O)N1CCC[C@@H]1CN1CCCC1
TAOK1TAOK1 over TAOK3
6.43predicted pIC50
v2 potency 0.5 of the reference ladder beaten · v2 prefers TAOK1 over TAOK3 at 0.718
Z2216912400
CN(c1cccc(CNc2nc(Nc3ccc4c(c3)CC(=O)N4)ncc2C(F)(F)F)c1)S(C)(=O)=O
CLK4CLK4 over CLK1
7.4predicted pIC50
v2 potency 1.0 of the reference ladder beaten · v2 prefers CLK4 over CLK1 at 0.714
Z44301537
CCOc1cc(C=C2SC(=S)NC2=O)ccc1O
LIMK1LIMK1 over LIMK2
6.42predicted pIC50
v2 potency 1.0 of the reference ladder beaten · v2 prefers LIMK1 over LIMK2 at 0.704
Z2612280147
Cc1[nH]c2ccccc2c1C1=CCN(c2cnc(N3CCCC3=O)cn2)CC1
CLK1CLK1 over CLK2
7.61predicted pIC50
v2 potency 0.688 of the reference ladder beaten · v2 prefers CLK1 over CLK2 at 0.785
Z89616206
COc1ccc2cc(CN(C)C(=O)c3cc4cc([N+](=O)[O-])ccc4s3)ccc2c1
CLK2CLK2 over CLK1exploratory
6.64predicted pIC50
v2 potency 0.812 of the reference ladder beaten · v2 prefers CLK2 over CLK1 at 0.614
Z7684579671
CC(C)CCOCCNc1nc(NC2COC(C3CC3)C2)c2ncn(C(C)C)c2n1
CLK1CLK1 over CLK4exploratory
7.02predicted pIC50
v2 potency 0.938 of the reference ladder beaten · v2 prefers CLK1 over CLK4 at 0.560
Z1918130497
Cc1cc(-c2cc(-c3nc(-c4cnc(N(C)C)cn4)no3)c3cnn(C(C)C)c3n2)c(C)o1

Predicted values are model output on compounds that have not been assayed. The survivor pool sits at median Tanimoto 0.27 to 0.36 from each target's training set, where measured external AUC is about 0.667: the ranking is informative, the absolute values are not.

The next step: designing what the catalogue does not contain

Everything above is selection. The three models narrowed 4,774,663 existing Enamine compounds down to twenty, and every slot in the panel is something you can already order. That is the right first move, because an orderable compound is the cheapest experiment available.

It is also a ceiling. A catalogue contains what someone already chose to make, so a selection campaign can only ever find the best answer that happens to exist. When the panel says a slot is exploratory, or when no catalogue compound sits where the models want one, the next step is to design the molecule instead.

ChIP does that without giving up makeability. It searches reaction space rather than molecule space: it evolves synthetic protocols built from validated reaction transforms and catalogued building blocks, then reads the molecule off the end of the route. Every design therefore arrives with an executable synthesis from purchasable starting materials, rather than a structure someone still has to work out how to make.

The two fit together directly, because the engine needs only one thing from a model: a number per molecule. Both comparators on this page supply one. The potency model scores a candidate by the fraction of a measured reference ladder it beats, exactly as it did in stage two above; the selectivity model scores the probability that a candidate prefers its target over a named sibling. Driving a ChIP campaign on the second puts the anti-target inside the objective function, which is what a selectivity claim requires and what a potency-only search cannot give you.

How ChIP works → A worked ChIP campaign →

Nothing in this panel has been assayed. Every figure above is a model prediction, and the compounds are real, orderable catalogue entries. Kinase Foundation Model · Eidogen-Sertanty · research software. Selection completed 11 August 2026 and reproduced here unchanged. Model limitations: /v2/limitations.html · KKB per-target models: eidogen-sertanty.com. Predictions are for research use and are not a substitute for measurement.