Eidogen-Sertanty · live results
Give it one of the 478 kinases it has data on, plus two molecules, and it says which one binds more tightly. That single comparison is the primitive operation of virtual screening, so repeating it sorts an entire library in seconds where a docking campaign takes hours — and at no point does the model see a three-dimensional structure. On molecules from a source it has never encountered it orders every target tested better than chance, and reaches 87 to 92 percent on decisive calls for its strongest targets. Which target you are asking about matters more than any other factor, and the reliability tiers below are the part to read before acting on a prediction.
A comparator. Given a kinase and descriptions of two drug-like molecules, it predicts which of the two binds more strongly. No three-dimensional structure, no docked pose and no binding site is required. Because a ranked library is nothing but this comparison repeated, the same model sorts a whole compound collection against a target in one pass. It was trained on human kinase measurements. The source corpus holds 841,187 rows over 500 genes; after the training split, embedding availability, removal of censored records and collapsing to one value per (sequence, ligand), 425,524 distinct measurements over 494 genes and 711 sequences remain eligible, and the forest is fitted on a random 250,000 of them, which covers fewer genes again.
Which kinases, precisely. The shipped model carries precomputed protein vectors for 478 named kinases over 695 distinct sequences, and those are the targets it will score. Being scoreable is not the same as being supported: the vectors were built for the whole panel, not for the subset the forest actually learned from, and the random 250,000-row fit covers fewer genes than the eligible pool. One exposed target, NEK10, has no usable training measurement anywhere in the corpus. Check a target against the reliability tiers before acting on any prediction for it, and treat a target absent from those tiers as unsupported. Handed a sequence it has never seen it refuses rather than guessing: a confident number derived from the wrong protein is the worst failure this system could produce. Nothing here supports a kinase with no training data.
Two compounds, one kinase, which binds harder. Coin flip = 50%.
Every prediction carries a confidence. Rank the calls by it and keep only the top slice: accuracy rises steeply as the slice narrows, so the operating point is yours to choose. The finest slice shown is the top tenth, which is the finest the confidence table resolves.
| calls acted on | pairs | accuracy |
|---|---|---|
| top 10% | 296,527 | 98.4% |
| top 25% | 741,318 | 95.7% |
| top 50% | 1,482,636 | 90.2% |
| top 75% | 2,223,955 | 83.9% |
| every pair | 2,965,273 | 77.2% |
| method | pairs | accuracy |
|---|---|---|
| Rank by molecular size alone | 2,807,614 | 58.5% |
| Use the compound's measured potency against a different kinase, where such a measurement exists | 1,492,766 | 73.1% |
| This model, on those same pairs | 1,492,766 | 76.8% |
| This model, where no prior measurement exists | 1,472,507 | 77.6% |
Size alone is a real bar and any useful model must clear it. The lookup row is the one that matters commercially: where a compound already has data against some other kinase, that lookup is a strong predictor and the model beats it only modestly. The model earns its place on the compounds where no such measurement exists, which is most of a screening library.
Evaluation is capped at 20,000 pairs per kinase so that a few very large targets cannot dominate; 119 kinases hit that cap. Reported, never silent.
The hardest test the dataset contains: pairs of compounds whose potency ordering reverses between two kinases. Compound A beats B on one target and loses to it on the other, both differences decisive. A model that cannot see the target produces one ordering for the pair and is therefore wrong on exactly one of the two kinases. It scores 50.0 percent by construction, not by chance. Anything above that line is target information and can be nothing else.
30,650 reversing cases drawn from 5,264 distinct compounds across 377 kinases. Reversals are 18.3 percent of the cases in this construction, which caps each gene pair at 60 shared ligands so a few heavily assayed pairs cannot dominate. It is not a kinome-wide reversal rate.
| model | accuracy on reversing pairs | Matthews | accuracy on non-reversing pairs |
|---|---|---|---|
| This model | 71.2% | 0.425 | 89.1% |
| The same forest, protein input removed | 50.0% | 0.000 | 89.0% |
The control lands exactly where arithmetic requires it to. The arm given no protein input at all scores 50.0 percent on the reversing pairs, confirming the test set is built correctly and that the margin above it is real. Shuffling the sequences rather than removing them gives 56.6 percent, a residue of signal from sequence length and amino acid composition and far below the model given the true sequence.
The flip analysis was run on the Morgan random forest, which differs from the published model only in fourteen size and composition descriptors whose measured contribution is +0.0023. The two score 0.7228 and 0.7232 on the main benchmark.
Screening needs a ranked list, not a single comparison. We gave the model a target and a whole set of compounds, had it sort them from most to least potent, and checked that order against what the assays actually measured.
We did this twice. First on compounds from our own collection that were held back from training. Then on compounds from ChEMBL, the public literature database, which come from different laboratories and which the model has never seen. The second is the harder test, and it is what the external numbers below describe. It is a retrospective benchmark on an outside dataset with exact training compounds removed, not an untouched independent replication: this corpus has been used elsewhere in the project, and only exact compound matches were excluded, not close analogues.
| same collection, compounds withheld | different database, never seen | |
|---|---|---|
| compounds ranked | 49,211 | 10,161 |
| pairwise comparisons | 46,330,109 | 1,926,222 |
| agreement with the measured order | 0.79 | 0.47 |
| pairs called right when potencies differ by more than tenfold | 90% | 77% |
| enrichment in the top tenth of the list | 5.5x | 3.6x |
Performance is not uniform and the aggregate hides that. These are the same 30 targets, split by how well the model ordered compounds it had never seen. A target's tier is the first thing to check before acting on a prediction for it.
Order these with confidence. Agreement of 0.6 or better, and the top of the list is strongly enriched in genuinely potent compounds.
| target | compounds | agreement with measured order | decisive pairs correct | enrichment in top tenth |
|---|---|---|---|---|
| AURKA | 221 | 0.83 | 92% | 2.7x |
| ABL1 | 496 | 0.77 | 87% | 3.8x |
| MTOR | 169 | 0.77 | 89% | 4.7x |
| SRC | 370 | 0.71 | 88% | 5.7x |
| CDK2 | 293 | 0.68 | 86% | 4.5x |
| PIM1 | 162 | 0.67 | 82% | 5.1x |
| BRAF | 217 | 0.62 | 83% | 4.5x |
Clearly better than chance and worth using to prioritise, but the ordering is loose enough that the top of the list should be confirmed rather than trusted.
| target | compounds | agreement with measured order | decisive pairs correct | enrichment in top tenth |
|---|---|---|---|---|
| CDK1 | 137 | 0.60 | 81% | 4.9x |
| FGFR1 | 477 | 0.59 | 82% | 3.1x |
| ERBB2 | 394 | 0.55 | 77% | 2.8x |
| MET | 467 | 0.50 | 79% | 4.7x |
| CDK9 | 442 | 0.47 | 77% | 3.9x |
| JAK3 | 377 | 0.46 | 73% | 5.2x |
| MAPK14 | 145 | 0.46 | 74% | 4.4x |
| SYK | 363 | 0.45 | 79% | 5.9x |
| PIK3CG | 454 | 0.43 | 72% | 3.6x |
| FLT1 | 132 | 0.43 | 73% | 4.7x |
| TYK2 | 487 | 0.42 | 71% | 3.4x |
Either the ordering is barely better than chance, or the top of the list is not enriched enough to act on — a target whose top tenth holds no more potent compounds than a random draw is not usable for screening however well the rest of the list is sorted. They are named rather than folded into an average.
| target | compounds | agreement with measured order | decisive pairs correct | enrichment in top tenth |
|---|---|---|---|---|
| AKT1 | 371 | 0.75 | 85% | 1.9x |
| PIK3CA | 371 | 0.59 | 78% | 1.9x |
| CDK4 | 214 | 0.50 | 82% | 1.0x |
| KDR | 274 | 0.46 | 74% | 1.1x |
| BTK | 398 | 0.40 | 71% | 2.5x |
| GSK3B | 332 | 0.35 | 67% | 2.4x |
| FLT3 | 467 | 0.33 | 65% | 3.6x |
| JAK2 | 258 | 0.33 | 65% | 4.6x |
| PDGFRB | 393 | 0.31 | 74% | 0.3x |
| MAPK1 | 499 | 0.29 | 65% | 1.0x |
| EGFR | 305 | 0.28 | 64% | 2.4x |
| PIK3CD | 476 | 0.09 | 56% | 2.5x |
The headline describes no individual target. Across 332 kinases with enough held-out pairs to measure, 237 reach 70% pairwise accuracy or better, 76 sit between 60% and 70%, and 19 fall below 60% against a 50% coin flip. Use it where it is strong; the rows below say where that is.
Arm: c46_rf_published. Sorted by the first metric column, best first — pairwise accuracy in ranker mode, rank correlation otherwise. Test ligands is the number of held-out compounds that target was scored on. Read it before the accuracy: a target scored on 50 pairs carries only about a dozen distinct ligands, and a high number there is not the same evidence as the same number over thousands of pairs.
| kinase | pairwise accuracy | MCC | test pairs | test ligands |
|---|---|---|---|---|
| ULK2 | 92.3% | 0.849 | 52 | 11 |
| EPHB1 | 90.8% | 0.817 | 65 | 12 |
| CAMK2D | 89.5% | 0.791 | 19730 | 279 |
| STK25 | 87.7% | 0.751 | 73 | 13 |
| DCLK1 | 87.4% | 0.743 | 166 | 19 |
| OXSR1 | 86.9% | 0.732 | 84 | 15 |
| CSNK1D | 85.6% | 0.713 | 19915 | 433 |
| RIPK1 | 85.2% | 0.703 | 19958 | 915 |
| MAP3K10 | 85.1% | 0.701 | 268 | 24 |
| IGF1R | 84.9% | 0.697 | 19781 | 694 |
| PIM1 | 84.6% | 0.693 | 19939 | 1689 |
| RIOK1 | 84.6% | 0.692 | 78 | 13 |
| EIF2AK2 | 84.6% | 0.682 | 65 | 12 |
| PTK2 | 84.5% | 0.691 | 19620 | 745 |
| RAF1 | 84.3% | 0.687 | 19880 | 834 |
| TTK | 84.2% | 0.683 | 19911 | 741 |
| MTOR | 84.0% | 0.680 | 19975 | 2222 |
| LTK | 83.9% | 0.676 | 1851 | 62 |
| PKMYT1 | 83.8% | 0.679 | 136 | 17 |
| STK33 | 83.8% | 0.678 | 117 | 16 |
| CAMK2A | 83.6% | 0.671 | 1302 | 52 |
| FGFR1 | 83.5% | 0.671 | 19962 | 1176 |
| ROCK1 | 83.4% | 0.667 | 19952 | 1002 |
| WEE1 | 83.2% | 0.664 | 19905 | 301 |
| TSSK1B | 83.0% | 0.659 | 393 | 29 |
| ABL1 | 82.9% | 0.659 | 19959 | 1145 |
| MST1R | 82.9% | 0.658 | 1806 | 61 |
| CDK4 | 82.9% | 0.658 | 19963 | 1374 |
| IKBKE | 82.8% | 0.657 | 19914 | 256 |
| MAPKAPK2 | 82.8% | 0.655 | 19951 | 413 |
| MAP4K1 | 82.5% | 0.649 | 19326 | 247 |
| AXL | 82.3% | 0.645 | 19889 | 468 |
| PRKCB | 82.2% | 0.644 | 19966 | 215 |
| CAMK1D | 82.2% | 0.642 | 689 | 38 |
| MAPK9 | 82.1% | 0.643 | 19906 | 402 |
| ALK | 82.1% | 0.642 | 19930 | 697 |
| PIK3CA | 82.1% | 0.641 | 19970 | 2066 |
| PRKD3 | 82.0% | 0.640 | 10463 | 147 |
| CDK2 | 82.0% | 0.640 | 19901 | 2211 |
| ABL2 | 82.0% | 0.639 | 560 | 34 |
| FGFR2 | 81.9% | 0.638 | 19917 | 777 |
| GRK5 | 81.9% | 0.637 | 574 | 35 |
| NTRK2 | 81.8% | 0.636 | 19918 | 213 |
| FGFR3 | 81.8% | 0.636 | 19952 | 910 |
| MAP3K9 | 81.8% | 0.635 | 274 | 24 |
| GSK3A | 81.7% | 0.634 | 19956 | 772 |
| ROS1 | 81.7% | 0.633 | 4922 | 100 |
| PAK1 | 81.6% | 0.633 | 10665 | 147 |
| PRKCE | 81.6% | 0.632 | 6420 | 114 |
| PTK6 | 81.6% | 0.631 | 934 | 44 |
| TBK1 | 81.6% | 0.631 | 19883 | 329 |
| JAK2 | 81.4% | 0.628 | 19976 | 3500 |
| CSF1R | 81.2% | 0.624 | 19913 | 603 |
| PIK3R4 | 81.1% | 0.621 | 19923 | 375 |
| PRKCA | 81.0% | 0.621 | 19955 | 256 |
| JAK1 | 81.0% | 0.620 | 19930 | 2951 |
| NEK4 | 81.0% | 0.620 | 1262 | 52 |
| PDGFRB | 80.9% | 0.618 | 19968 | 636 |
| DMPK | 80.9% | 0.619 | 68 | 13 |
| ACVRL1 | 80.9% | 0.618 | 701 | 38 |
| SYK | 80.8% | 0.617 | 19712 | 1497 |
| PKN1 | 80.8% | 0.612 | 187 | 20 |
| EPHA6 | 80.6% | 0.611 | 252 | 23 |
| CSNK2A1 | 80.5% | 0.610 | 19956 | 449 |
| MAP2K2 | 80.5% | 0.610 | 1209 | 50 |
| BRAF | 80.5% | 0.610 | 19917 | 1455 |
| EGFR | 80.4% | 0.609 | 19959 | 2441 |
| PLK2 | 80.4% | 0.608 | 1319 | 52 |
| TYK2 | 80.3% | 0.607 | 19982 | 1259 |
| PIK3CG | 80.3% | 0.605 | 19966 | 1394 |
| AKT1 | 80.3% | 0.605 | 19957 | 993 |
| EPHA2 | 80.3% | 0.605 | 16295 | 187 |
| AURKC | 80.2% | 0.605 | 592 | 35 |
| PIK3CB | 80.1% | 0.603 | 19965 | 829 |
| MAP4K2 | 80.1% | 0.603 | 8992 | 136 |
| MAPK8 | 80.1% | 0.602 | 19948 | 635 |
| TYRO3 | 80.0% | 0.601 | 19896 | 252 |
| GRK7 | 80.0% | 0.601 | 90 | 14 |
| MELK | 80.0% | 0.600 | 19858 | 252 |
| SRC | 79.9% | 0.599 | 19960 | 1112 |
| PRKCQ | 79.9% | 0.598 | 19972 | 681 |
| CLK4 | 79.9% | 0.597 | 19836 | 353 |
| CLK1 | 79.9% | 0.597 | 19959 | 268 |
| STK24 | 79.8% | 0.595 | 104 | 15 |
| TNK2 | 79.6% | 0.593 | 12050 | 156 |
| DCLK2 | 79.6% | 0.594 | 265 | 24 |
| INSR | 79.6% | 0.592 | 19804 | 301 |
| CAMK2G | 79.6% | 0.591 | 3884 | 90 |
| RPS6KA3 | 79.5% | 0.591 | 19714 | 231 |
| EPHB2 | 79.5% | 0.592 | 273 | 24 |
| EPHA1 | 79.5% | 0.589 | 297 | 25 |
| GSK3B | 79.5% | 0.589 | 19946 | 1383 |
| CDK3 | 79.3% | 0.585 | 666 | 37 |
| AKT3 | 79.2% | 0.585 | 19909 | 317 |
| EPHA8 | 79.2% | 0.583 | 461 | 31 |
| NTRK1 | 79.2% | 0.583 | 19966 | 982 |
| PIM2 | 79.0% | 0.581 | 19964 | 1117 |
| TEC | 79.0% | 0.580 | 816 | 41 |
| FLT4 | 79.0% | 0.580 | 19871 | 216 |
| MINK1 | 78.9% | 0.577 | 4977 | 102 |
| HCK | 78.9% | 0.577 | 12800 | 161 |
| KDR | 78.8% | 0.577 | 19934 | 3026 |
| TAOK3 | 78.8% | 0.576 | 453 | 32 |
| HIPK1 | 78.7% | 0.574 | 376 | 28 |
| CSNK2A2 | 78.6% | 0.573 | 4643 | 97 |
| NTRK3 | 78.5% | 0.571 | 10342 | 145 |
| AURKB | 78.5% | 0.569 | 19908 | 876 |
| LRRK2 | 78.5% | 0.569 | 19965 | 1048 |
| TXK | 78.4% | 0.570 | 662 | 37 |
| MAPK10 | 78.4% | 0.568 | 19939 | 562 |
| IRAK4 | 78.4% | 0.567 | 19940 | 1020 |
| PRKAA1 | 78.3% | 0.567 | 12487 | 159 |
| STK4 | 78.2% | 0.569 | 170 | 19 |
| SIK1 | 78.2% | 0.565 | 19981 | 231 |
| MAP4K3 | 78.1% | 0.565 | 119 | 16 |
| MAPK12 | 78.1% | 0.562 | 4249 | 93 |
| PRKCD | 78.0% | 0.560 | 19899 | 294 |
| AKT2 | 77.9% | 0.559 | 19916 | 329 |
| FGR | 77.9% | 0.557 | 3641 | 86 |
| PTK2B | 77.9% | 0.558 | 7247 | 121 |
| DYRK1A | 77.9% | 0.558 | 19930 | 660 |
| CDK1 | 77.9% | 0.557 | 19957 | 1008 |
| PDGFRA | 77.8% | 0.556 | 19938 | 432 |
| FGFR4 | 77.8% | 0.555 | 19951 | 606 |
| MAPK1 | 77.7% | 0.555 | 19976 | 1110 |
| PLK4 | 77.6% | 0.553 | 19931 | 255 |
| MAP4K4 | 77.6% | 0.552 | 19713 | 276 |
| PRKDC | 77.5% | 0.550 | 19860 | 234 |
| PRKCH | 77.4% | 0.548 | 5020 | 101 |
| ATM | 77.3% | 0.546 | 5619 | 107 |
| PDK2 | 77.2% | 0.544 | 11422 | 152 |
| PBK | 77.2% | 0.544 | 5317 | 104 |
| CSNK1E | 77.2% | 0.544 | 19968 | 276 |
| MAP3K11 | 77.2% | 0.542 | 206 | 21 |
| PRKCG | 77.1% | 0.542 | 4848 | 100 |
| JAK3 | 77.1% | 0.541 | 19973 | 1950 |
| MERTK | 77.0% | 0.541 | 19918 | 310 |
| MARK1 | 77.0% | 0.534 | 87 | 14 |
| AURKA | 77.0% | 0.539 | 19959 | 1247 |
| CDC42BPG | 76.9% | 0.535 | 65 | 12 |
| ERBB4 | 76.8% | 0.536 | 5840 | 109 |
| MAPK3 | 76.8% | 0.542 | 250 | 23 |
| LIMK1 | 76.8% | 0.536 | 19850 | 202 |
| ACVR1 | 76.8% | 0.536 | 7828 | 126 |
| MET | 76.7% | 0.535 | 19940 | 1211 |
| RPS6KB1 | 76.7% | 0.534 | 19956 | 664 |
| MAPK15 | 76.7% | 0.530 | 120 | 16 |
| PIK3CD | 76.6% | 0.532 | 19860 | 1225 |
| MAPK14 | 76.5% | 0.531 | 19976 | 1862 |
| CSK | 76.5% | 0.528 | 481 | 32 |
| SRMS | 76.5% | 0.529 | 978 | 45 |
| ERBB2 | 76.4% | 0.529 | 19905 | 1330 |
| FLT1 | 76.4% | 0.528 | 19889 | 1128 |
| MKNK2 | 76.4% | 0.527 | 19919 | 399 |
| LCK | 76.3% | 0.527 | 19928 | 624 |
| DAPK1 | 76.3% | 0.524 | 135 | 17 |
| CLK2 | 76.3% | 0.525 | 19888 | 313 |
| YES1 | 76.3% | 0.525 | 4262 | 93 |
| FLT3 | 76.2% | 0.525 | 19956 | 1152 |
| PRKX | 76.2% | 0.524 | 3554 | 86 |
| CHEK1 | 76.1% | 0.523 | 19949 | 639 |
| CDK5 | 76.0% | 0.520 | 19913 | 507 |
| MARK3 | 76.0% | 0.519 | 3463 | 85 |
| PLK1 | 75.9% | 0.519 | 19942 | 341 |
| PDK4 | 75.7% | 0.514 | 136 | 17 |
| KIT | 75.7% | 0.514 | 19906 | 657 |
| BTK | 75.6% | 0.511 | 19961 | 1608 |
| BRSK1 | 75.6% | 0.511 | 1522 | 57 |
| EPHB4 | 75.5% | 0.510 | 16707 | 187 |
| PRKCZ | 75.4% | 0.509 | 1693 | 59 |
| MAPK13 | 75.3% | 0.506 | 1097 | 48 |
| SIK3 | 75.2% | 0.504 | 19949 | 219 |
| PAK4 | 75.1% | 0.503 | 13603 | 166 |
| MKNK1 | 75.1% | 0.503 | 19950 | 371 |
| CDK6 | 75.0% | 0.501 | 19960 | 526 |
| RIOK2 | 75.0% | 0.501 | 64 | 12 |
| CDK19 | 75.0% | 0.499 | 4720 | 98 |
| DYRK1B | 74.8% | 0.496 | 19911 | 271 |
| IKBKB | 74.7% | 0.495 | 19934 | 405 |
| PDPK1 | 74.7% | 0.493 | 19914 | 361 |
| PKN2 | 74.6% | 0.492 | 3178 | 81 |
| DDR1 | 74.6% | 0.492 | 2478 | 71 |
| MAP2K1 | 74.6% | 0.491 | 19955 | 340 |
| TGFBR1 | 74.5% | 0.491 | 19963 | 593 |
| ROCK2 | 74.5% | 0.490 | 19971 | 961 |
| MAP3K3 | 74.4% | 0.486 | 78 | 13 |
| BLK | 74.3% | 0.487 | 5580 | 107 |
| MAP2K6 | 74.3% | 0.497 | 113 | 16 |
| RPS6KA1 | 74.3% | 0.486 | 19920 | 416 |
| IRAK3 | 74.3% | 0.486 | 105 | 15 |
| MAP3K12 | 74.3% | 0.486 | 15007 | 174 |
| AAK1 | 74.2% | 0.484 | 19426 | 198 |
| PRKD1 | 74.1% | 0.484 | 348 | 27 |
| PIM3 | 74.0% | 0.480 | 19796 | 977 |
| TAOK1 | 73.9% | 0.478 | 5759 | 110 |
| TEK | 73.9% | 0.477 | 16436 | 182 |
| MAP3K14 | 73.9% | 0.477 | 19907 | 281 |
| ITK | 73.8% | 0.476 | 19930 | 444 |
| PRKG2 | 73.7% | 0.475 | 875 | 43 |
| PRKCI | 73.7% | 0.474 | 9591 | 160 |
| GRK2 | 73.6% | 0.473 | 91 | 14 |
| RET | 73.5% | 0.471 | 19982 | 770 |
| LYN | 73.4% | 0.469 | 19905 | 205 |
| NEK2 | 73.4% | 0.468 | 3577 | 86 |
| MAP3K7 | 73.3% | 0.466 | 8545 | 132 |
| PRKG1 | 73.0% | 0.460 | 1599 | 58 |
| SIK2 | 72.9% | 0.457 | 19949 | 331 |
| DAPK3 | 72.9% | 0.457 | 7159 | 121 |
| STK10 | 72.9% | 0.459 | 700 | 38 |
| CDK7 | 72.7% | 0.454 | 19822 | 474 |
| INSRR | 72.3% | 0.430 | 130 | 17 |
| BMPR1A | 72.3% | 0.446 | 231 | 22 |
| PDK1 | 72.3% | 0.446 | 16035 | 248 |
| CAMK1 | 72.3% | 0.445 | 249 | 23 |
| TSSK2 | 72.2% | 0.415 | 54 | 11 |
| ACVR1B | 72.1% | 0.439 | 208 | 21 |
| CSNK1G1 | 72.1% | 0.441 | 1354 | 53 |
| MARK2 | 72.0% | 0.441 | 3700 | 88 |
| PI4KB | 72.0% | 0.440 | 5022 | 101 |
| FYN | 71.9% | 0.438 | 19958 | 342 |
| DDR2 | 71.9% | 0.438 | 1427 | 54 |
| PRKACA | 71.7% | 0.434 | 19938 | 281 |
| SGK1 | 71.7% | 0.434 | 16695 | 184 |
| CSNK1G2 | 71.6% | 0.432 | 2682 | 75 |
| ULK1 | 71.6% | 0.433 | 88 | 14 |
| IRAK1 | 71.5% | 0.430 | 11915 | 156 |
| CDK9 | 71.4% | 0.428 | 19931 | 1322 |
| PIK3C3 | 71.3% | 0.426 | 9287 | 137 |
| HIPK2 | 71.0% | 0.421 | 4735 | 99 |
| CAMKK2 | 71.0% | 0.420 | 806 | 41 |
| RPS6KA2 | 71.0% | 0.419 | 1704 | 59 |
| CHEK2 | 70.5% | 0.409 | 13410 | 165 |
| EIF2AK4 | 70.4% | 0.408 | 54 | 11 |
| EPHA7 | 70.4% | 0.427 | 135 | 17 |
| CDK8 | 70.3% | 0.407 | 19883 | 248 |
| MAPKAPK5 | 70.2% | 0.404 | 1086 | 48 |
| MAP2K5 | 70.0% | 0.400 | 300 | 25 |
| SGK2 | 69.9% | 0.398 | 1283 | 52 |
| PRKD2 | 69.7% | 0.395 | 6530 | 116 |
| CDC42BPA | 69.6% | 0.392 | 2641 | 75 |
| MYLK | 69.6% | 0.391 | 450 | 31 |
| RIPK2 | 69.5% | 0.391 | 11239 | 151 |
| CLK3 | 69.5% | 0.390 | 4269 | 93 |
| PIP5K1C | 69.2% | 0.323 | 52 | 11 |
| TNK1 | 69.2% | 0.363 | 91 | 14 |
| BMPR2 | 69.2% | 0.391 | 91 | 14 |
| EPHA5 | 69.2% | 0.391 | 120 | 16 |
| ERBB3 | 69.1% | 0.382 | 55 | 11 |
| DGKZ | 69.0% | 0.381 | 19891 | 270 |
| EPHB6 | 69.0% | 0.374 | 190 | 20 |
| MYLK2 | 68.9% | 0.378 | 90 | 14 |
| CDC7 | 68.8% | 0.376 | 19923 | 204 |
| ATR | 68.6% | 0.373 | 9021 | 135 |
| ERN1 | 68.6% | 0.367 | 153 | 18 |
| CAMK4 | 68.6% | 0.381 | 51 | 12 |
| MAP3K8 | 68.6% | 0.372 | 937 | 44 |
| MARK4 | 68.4% | 0.369 | 76 | 13 |
| STK17B | 68.2% | 0.379 | 88 | 14 |
| EPHA3 | 68.2% | 0.338 | 66 | 12 |
| NEK1 | 68.2% | 0.362 | 135 | 17 |
| MAP3K2 | 68.1% | 0.363 | 348 | 27 |
| PHKG2 | 68.1% | 0.361 | 1040 | 47 |
| DYRK2 | 68.1% | 0.361 | 14159 | 169 |
| HIPK3 | 68.0% | 0.356 | 231 | 22 |
| DAPK2 | 68.0% | 0.359 | 206 | 21 |
| STK26 | 67.8% | 0.364 | 90 | 14 |
| IKBKG | 67.5% | 0.349 | 3790 | 88 |
| STK17A | 67.4% | 0.347 | 3098 | 80 |
| TNIK | 67.4% | 0.348 | 1428 | 54 |
| MAPKAPK3 | 67.3% | 0.346 | 165 | 19 |
| DYRK4 | 67.1% | 0.341 | 2829 | 76 |
| CSNK1A1 | 66.7% | 0.334 | 19070 | 197 |
| MAP3K19 | 66.7% | 0.330 | 135 | 17 |
| STK3 | 66.6% | 0.331 | 8684 | 133 |
| MAP3K20 | 66.5% | 0.330 | 2327 | 70 |
| EPHA4 | 66.5% | 0.320 | 152 | 18 |
| MAP2K7 | 66.4% | 0.328 | 298 | 25 |
| NUAK2 | 66.4% | 0.328 | 134 | 17 |
| CDK11A | 66.3% | 0.324 | 89 | 14 |
| LIMK2 | 66.2% | 0.323 | 1826 | 61 |
| PAK2 | 66.1% | 0.323 | 3823 | 89 |
| NUAK1 | 65.8% | 0.306 | 120 | 16 |
| DYRK3 | 65.6% | 0.311 | 8324 | 130 |
| CDK13 | 65.6% | 0.315 | 90 | 14 |
| ULK3 | 65.6% | 0.312 | 90 | 14 |
| TAOK2 | 65.4% | 0.316 | 133 | 17 |
| CHUK | 65.3% | 0.307 | 7040 | 120 |
| STK16 | 65.3% | 0.309 | 300 | 25 |
| RPS6KA5 | 64.9% | 0.299 | 1959 | 64 |
| FES | 64.8% | 0.298 | 523 | 33 |
| BUB1 | 64.6% | 0.292 | 19893 | 233 |
| GSG2 | 64.6% | 0.295 | 325 | 26 |
| ZAP70 | 64.6% | 0.292 | 1940 | 63 |
| CAMK2B | 64.3% | 0.286 | 2042 | 67 |
| BMPR1B | 64.3% | 0.286 | 350 | 27 |
| CSNK2B | 64.3% | 0.285 | 750 | 40 |
| CDK12 | 64.2% | 0.284 | 7264 | 131 |
| DGKA | 64.2% | 0.283 | 19851 | 339 |
| FRK | 64.1% | 0.282 | 19784 | 298 |
| RPS6KA6 | 64.0% | 0.280 | 297 | 25 |
| MAP3K13 | 63.6% | 0.187 | 66 | 12 |
| MAP2K4 | 63.6% | 0.260 | 66 | 12 |
| MAPK11 | 63.4% | 0.266 | 3223 | 81 |
| CDK16 | 63.1% | 0.263 | 350 | 27 |
| CSNK1G3 | 63.1% | 0.262 | 837 | 42 |
| HIPK4 | 63.0% | 0.261 | 2748 | 76 |
| PLK3 | 62.7% | 0.254 | 3674 | 87 |
| EIF2AK3 | 62.6% | 0.259 | 147 | 18 |
| CDK17 | 62.5% | 0.293 | 64 | 12 |
| GRK6 | 61.5% | 0.281 | 65 | 12 |
| MAP4K5 | 61.3% | 0.225 | 8900 | 136 |
| MATK | 60.3% | 0.205 | 219 | 22 |
| SLK | 60.1% | 0.202 | 4635 | 98 |
| TNNI3K | 59.5% | 0.190 | 887 | 43 |
| GAK | 59.2% | 0.185 | 1325 | 52 |
| FER | 59.0% | 0.180 | 1049 | 47 |
| CDKL2 | 58.5% | 0.159 | 65 | 12 |
| BMP2K | 57.9% | 0.160 | 349 | 27 |
| NLK | 57.8% | 0.155 | 187 | 20 |
| BMX | 57.0% | 0.139 | 4308 | 98 |
| GRK1 | 56.6% | 0.119 | 76 | 13 |
| TIE1 | 55.8% | 0.112 | 77 | 13 |
| RPS6KA4 | 55.6% | 0.119 | 189 | 20 |
| SRPK1 | 55.4% | 0.108 | 1392 | 55 |
| TGFBR2 | 54.6% | 0.094 | 346 | 27 |
| CIT | 53.6% | 0.070 | 209 | 21 |
| RIPK3 | 52.9% | 0.064 | 119 | 16 |
| CSNK2A3 | 49.1% | -0.030 | 118 | 16 |
| MAP2K3 | 48.7% | -0.025 | 119 | 16 |
| MAPK7 | 47.3% | -0.092 | 55 | 11 |
| PIP4K2C | 47.0% | -0.061 | 66 | 12 |
| ICK | 39.5% | -0.221 | 119 | 16 |
Of the 179 targets this project classes as understudied, the published model already scores 163, because they are targets with modest data rather than targets with none. Taken at face value they look weaker: median accuracy 0.736 against 0.769 for the rest.
They also carry far less held-out data — a median of 754 pairs against 19,893, a twenty-six-fold difference. Comparing them at matched data volume, in the one band where both groups are well populated, understudied targets score 0.715 (n=31) against 0.723 (n=22): a difference of -0.0077 ± 0.0403, indistinguishable from zero.
The model answers a comparison: given one kinase sequence and two ligands, which ligand is more potent against that kinase. Ranking a compound library is a sequence of such comparisons, so this is the primitive operation of virtual screening rather than a proxy for it. Because the orientation of every pair is randomised, a constant output scores exactly 0.500 and a Matthews correlation of 0.000, so the chance line needs no argument. That is not the same as there being no baseline worth beating, and two are reported beside the model below: ranking by molecular size alone, and looking a compound's potency up from another kinase where such data exists.
Measurements are drawn from the Kinase Knowledgebase (KKB), Eidogen-Sertanty, release Q2-2026, restricted to human enzyme assays with a defined potency between pIC50 3 and 11. The training table holds 673,659 measurements over 500 kinase genes and 742 distinct sequences, including 13,218 rows carrying a mutated sequence; a mutant is represented by its own sequence rather than being folded into the wild type. Censored records, which state only that potency is beyond a bound, are excluded from pair formation because a bound cannot be ordered against an exact value.
After excluding censored records and collapsing to one measurement per (sequence, ligand) pair, 417,507 distinct measurements remain. The published model is fitted on a random sample of 250,000 of these, about 60 percent, because the dense feature matrix over all of them does not fit comfortably in memory. The sample is drawn once with a fixed seed and the cap is recorded in the run's result file. The model therefore reaches the accuracies reported here having seen roughly three fifths of the available training measurements.
Each ligand is described by a Morgan count fingerprint of radius 2 and 1,024 bits, computed with RDKit from the canonical SMILES, concatenated with fourteen physicochemical and compositional descriptors: molecular weight, heavy-atom count, bond count, rotatable bonds, ring count, and the counts of C, N, O, S, F, Cl, Br, I and P. The descriptors are part of the arm that was scored and are reported for exactness. Their measured contribution is negligible: paired per kinase over the 213 targets with at least 1,000 pairs, adding them is worth +0.0023 ± 0.0030 on count fingerprints. An earlier version of this document explained them by saying a circular fingerprint represents molecular size poorly; that explanation predicted a material gain on BINARY fingerprints, and the measured gain there was +0.0032 ± 0.0031, no different. The explanation is withdrawn and plain Morgan counts would serve as well.
Each kinase sequence is embedded with ESM2 (esm2_t12_35M_UR50D), giving one vector per residue, and reduced to a single fixed 480-dimensional vector by mean pooling over residues. Pooling is fixed rather than learned. Concatenated with the 1,038 ligand features, the vector presented to the model is 1,518 dimensions: 480 protein and 1,038 ligand.
The model is trained on human kinase sequences and ligand structures from the Kinase Knowledgebase, and on nothing else. Its training table carries 673,659 rows over 500 kinases and 742 distinct sequences, of which 13,218 rows are point-mutant sequence records. Every row is an original measurement.
The structure-free framing of this project — predicting engagement from protein sequence and ligand chemistry alone, with no three-dimensional pose, using a protein language model together with a chemical language model — follows Fondrie and colleagues, Structure-free, site-resolved contrastive learning extends small-molecule discovery beyond the reach of structure-based modeling (Talus Bioscience, bioRxiv 2026), which introduces Ptarmigan-1.
The debt is specific and worth stating precisely. This project's neural arm was adapted from that architecture: frozen ESM2 residue embeddings and frozen ChemBERTa ligand embeddings projected into a shared space, with the temperature-scaled softmax pooling over residues that Ptarmigan-1 uses to turn residue-level scores into a protein-level call reused here as attention pooling, and its leakage-safe Bemis-Murcko scaffold-split discipline. Trainable projections and head only, both backbones frozen, as a lighter-weight analogue of its LoRA fine-tuning.
The model reported here is not that model. The neural arm was beaten on every measure by the random forest described above, which uses fixed mean pooling and a count fingerprint rather than learned pooling and a chemical language model, and which is what all results in this document describe. The two also answer different questions: Ptarmigan-1 predicts and localises engagement across the proteome, including at cryptic and disordered sites; this model ranks ligands by potency against a kinase for which measured data already exists.
A random forest regressor of 200 trees (minimum two samples per leaf) is fitted pointwise on (protein vector, ligand features) against measured pIC50. Two ligands are then compared by the difference of the model's two predictions, so that P(A more potent than B) follows the sign of sA − sB. This construction is exactly antisymmetric: exchanging the two ligands reverses the prediction identically, and the model cannot return contradictory answers for the same pair presented in a different order. A forest given both ligands as joint input would carry no such guarantee. It also means a library of n compounds is scored in n forward passes rather than n2 comparisons.
This model has no training pairs. The forest is fitted pointwise on individual pIC50 values; pairs exist only at evaluation, and every non-tied pair of ligands measured against the same sequence is scored. There is no gap filter on pair construction.
Results are then reported in three bands of true potency difference: under half a log, half to one log, and more than one log. The band most worth quoting is the last, because that is the separation a chemist acts on. For context on what a resolvable difference is, disagreement between two independent publications reporting the same compound against the same kinase has a 90th percentile of 0.715 log units over 204,438 comparisons, rising to 1.218 after deduplication; the source file recommends treating one log as the floor. Pairs closer than that are harder rather than impossible: the model still scores 57.8 percent under half a log, against a 50 percent coin flip.
The advantage that cross-kinase knowledge confers is not assumed away, it is measured. The compound-lookup baseline above is exactly that advantage made explicit, predicting from the same compound’s potency against another kinase, and it scores 73.1 percent. The model scores 76.8 percent on those same pairs, so it adds information beyond knowing the compound elsewhere. What these internal numbers do not measure is performance on chemistry absent from the corpus entirely; the external section speaks to that.
Ligands are withheld by Bemis-Murcko scaffold, keyed on the (sequence, ligand) pair. Every kinase sequence remains in training. No protein and no sequence family is withheld, because a kinase for which no data exists is not the deployment scenario: the kinome is among the most exhaustively characterised regions of the proteome. Evaluation covers 2,965,273 held-out ligand pairs across 413 targets. The kinase is the unit of replication for the kinase-level mean and its standard error, and only for those. Pooled accuracy, the margin deciles and the selectivity figures are computed over pairs that share compounds and are therefore not independent observations; treat them as descriptive rather than as quantities with an interval.
The headline results are obtained without three bodies of work that are built and available: 4,555 CANDIDATE orthologue sequences across 475 organisms, screened by binding-site identity; per-residue contact maps from 1,417 solved co-complex structures covering 170 targets, with 977 identified ligands and 182 sites excluded because a nucleotide or crystallisation additive rather than an inhibitor defined them; and 821,370 BindingDB measurements retained after removing every compound-target pair appearing in the frozen ChEMBL benchmarks. These are headroom, not caveats.
Kinase Knowledgebase (KKB), Eidogen-Sertanty, release Q2-2026 — the training and evaluation corpus: eidogen-sertanty.com/kinasekbmarvin.php.
BindingDB — ingested and held for future arms, not used in the results reported here. Gilson et al., Nucleic Acids Res. 44:D1045 (2016), doi:10.1093/nar/gkv1072; bindingdb.org.
ChEMBL — used only as a frozen external benchmark, never trained on; every compound-target pair appearing in it was removed from the BindingDB import. Zdrazil et al., Nucleic Acids Res. 52:D1180 (2024), doi:10.1093/nar/gkad1004; ebi.ac.uk/chembl.
RCSB Protein Data Bank — co-complex structures used for the binding-site analysis reported separately. Berman et al., Nucleic Acids Res. 28:235 (2000), doi:10.1093/nar/28.1.235; rcsb.org.
UniProt — sequence and accession mapping. UniProt Consortium, Nucleic Acids Res. 51:D523 (2023), doi:10.1093/nar/gkac1052; uniprot.org.
ESM2 protein language model, checkpoint facebook/esm2_t12_35M_UR50D, 35M parameters, 12 layers, 480-dimensional per-residue representations. Lin et al., Science 379:1123 (2023), doi:10.1126/science.ade2574; code at github.com/facebookresearch/esm.
ChemBERTa chemical language model, checkpoint seyonec/ChemBERTa-zinc-base-v1, 768-dimensional embeddings — used in the comparison arm. Chithrananda et al., arXiv:2010.09885 (2020), arxiv.org/abs/2010.09885.
Morgan / ECFP count fingerprints, radius 2, 1,024 bits, and all physicochemical descriptors computed with RDKit: rdkit.org. Method: Rogers & Hahn, J. Chem. Inf. Model. 50:742 (2010), doi:10.1021/ci100050t.
Bemis-Murcko scaffolds, used to define the held-out ligand split. Bemis & Murcko, J. Med. Chem. 39:2887 (1996), doi:10.1021/jm9602928.
Random forest regressor, scikit-learn: scikit-learn.org. Breiman, Machine Learning 45:5 (2001), doi:10.1023/A:1010933404324.
Pairwise ranking formulation: Burges et al., Learning to Rank using Gradient Descent, ICML 2005, doi:10.1145/1102351.1102363.
Matthews correlation coefficient: Matthews, Biochim. Biophys. Acta 405:442 (1975), doi:10.1016/0005-2795(75)90109-9.
Point it at a kinase you already have data on and a compound library you have not measured. It scores every compound in one pass and returns a ranked list, each comparison carrying a confidence. Act on the confident fraction: accuracy rises steeply as you narrow to the calls the model is most certain about, so the operating point is yours to choose.
Check the target first. This is the single most important step and it is not optional. Performance varies more between targets than it does between any two models we have tried. On its strongest targets — AURKA, ABL1, MTOR, SRC, CDK2, PIM1, BRAF — it calls decisive pairs correctly around nine times in ten on chemistry from a source it has never seen. On its weakest it is close enough to a coin flip that a prediction should carry no weight. The tier table above is the guide, and there is no aggregate number that substitutes for reading it.
It ranks; it does not predict absolute potency. Two compounds whose true potencies differ by less than two publications routinely disagree by cannot be reliably adjudicated from this assay corpus, and the model is correctly close to chance on them rather than confidently wrong.
Where a compound already has measured potency against a related kinase, use that measurement: it is direct evidence. It is not, however, a stronger predictor than the model on this benchmark — that lookup scores 73.1 percent where it is defined, against the model's 76.8 percent on the same pairs. The lookup was undefined on 49.7 percent of evaluation pairs, and it is there that the model is the only option.
It is a triage tool that runs before docking or crystallography, not a replacement for either, and not a substitute for measurement.
It sorts a compound library against a kinase you already have data on, and reports a margin on each call that tracks how often it is right. Used as a filter on its confident predictions it is right about nine times in ten on chemistry from a source it has never seen. A coin flip is 50%.
Scores a pair in milliseconds with no structure needed, against minutes to hours for docking, so it triages a library before any structure-based work begins. Structural evidence behind the binding-site work.
The complete experiment record, including every claim that did not
survive review, is kept separately in collab/cycles/ and is deliberately not
part of this document.
Generated from collab/cycles/*/claude_manifest.json by scripts/37_live_report.py · 08 August 2026 at 19:01