Eidogen-Sertanty · live results
Give it one of the 478 kinases it has data on, plus two molecules, and it says which one binds more tightly. That single comparison is the primitive operation of virtual screening, so repeating it sorts an entire library in seconds where a docking campaign takes hours, and at no point does the model see a three-dimensional structure.
Two models are released. The Validated model carries every number in this report, measured on 2,965,273 pairs it never saw. The Frontier model uses the identical training method on every measurement we hold, including the data withheld to make that measurement possible, and is therefore untested by construction. Which target you ask about matters more than which model you pick: the reliability tiers below are the part to read before acting on any prediction.
A comparator. Given a kinase and descriptions of two drug-like molecules, it predicts which of the two binds more strongly. No three-dimensional structure, no docked pose and no binding site is required. Because a ranked library is nothing but this comparison repeated, the same model sorts a whole compound collection against a target in one pass. It was trained on human kinase measurements. The source corpus holds 841,187 rows over 500 genes. One third is held out for testing; of the remainder, records giving only a bound rather than a value are set aside and repeat measurements of the same compound against the same kinase are collapsed to their median, leaving 417,507 distinct measurements. The model is fitted on 417,507 of them, all of them.
Which kinases, precisely. The shipped model carries precomputed protein vectors for 478 named kinases over 695 distinct sequences, and those are the targets it will score. Being scoreable is not the same as being supported: the vectors were built for the whole panel, not for the subset the forest actually learned from, so the fit covers only the genes that carry usable measurements. Three exposed targets, NEK10, ADK and SPHK2, contributed no fitted rows; NEK10 is the fundamental case, because every one of its measurements is a limit rather than an exact value. Check a target against the reliability tiers before acting on any prediction for it, and treat a target absent from those tiers as unsupported. Handed a sequence it has never seen it refuses rather than guessing: a confident number derived from the wrong protein is the worst failure this system could produce. Nothing here supports a kinase with no training data.
Two compounds, one kinase, which binds harder. Coin flip = 50%.
Every prediction carries a confidence. Rank the calls by it and keep only the top slice: accuracy rises steeply as the slice narrows, so the operating point is yours to choose. The finest slice shown is the top tenth, which is the finest the confidence table resolves.
| calls acted on | pairs | accuracy |
|---|---|---|
| top 10% | 296,527 | 98.8% |
| top 25% | 741,318 | 96.6% |
| top 50% | 1,482,636 | 91.8% |
| top 75% | 2,223,955 | 85.7% |
| every pair | 2,965,273 | 78.8% |
Two cheap shortcuts a screening group could use instead. Neither is an input to the model and neither changes what the model is given; they are rival predictors, and the model has to be better than both to be worth running.
| method | pairs | accuracy |
|---|---|---|
| Rank by molecular size alone | 2,807,614 | 58.5% |
| Shortcut 2: look up what this compound did against some other kinase and assume it behaves the same here | 1,492,766 | 73.1% |
| This model, on those same pairs | 1,492,766 | 78.7% |
| This model, where no prior measurement exists | 1,472,507 | 79.0% |
Shortcut 2 is the serious rival. Potency is correlated across kinases, so knowing what a compound did elsewhere predicts a lot, 73.1 percent, without any model at all. The model beats it on the same pairs, 76.8 percent, and is the only option on the 1.47 million pairs where no such prior measurement exists. Neither shortcut can address selectivity: both give one answer per compound regardless of which kinase you ask about.
Evaluation is capped at 20,000 pairs per kinase so that a few very large targets cannot dominate; 119 kinases hit that cap. Reported, never silent.
A fair worry about any model like this: maybe it never really uses the protein, and is just deciding which of the two compounds looks better in general. Here is the test that settles it.
Find compound pairs that were measured against two different kinases, and keep the ones where the answer switches:
A model that ignores the protein produces one answer for that pair. It is therefore right on one kinase and wrong on the other, every time. It scores 50 percent because arithmetic forces it to, not because it is guessing. To do better than 50 percent here, a model has no choice but to use the kinase.
| model | gets these switched cases right |
|---|---|
| This model | 71.2% |
| The identical model with the protein taken away | 50.0% |
30,650 switched cases, 5,264 compounds, 377 kinases. Both rows are the same random forest on the same pairs. The only difference is whether the protein vector changes from one kinase to the next.
Package identity. Every number in this document was produced by the forest shipped as kfm_ranker_v1_final, SHA-256 4f36bf27f5d7c0a7…. Reloading that exact file and rescoring reproduces the internal benchmark with a maximum per-kinase difference of 0.000000.
This has nothing to do with swapping the two compounds at the input. If you feed the same model compound B first and compound A second, the answer flips exactly, every time, by construction. That is a property of how the model is built and it is asserted in the package self-test. The test above is about something else entirely: the same pair giving genuinely opposite answers on two different proteins.
Screening needs a ranked list, not a single comparison. We gave the model a target and a whole set of compounds, had it sort them from most to least potent, and checked that order against what the assays actually measured.
We did this twice. First on compounds from our own collection that were held back from training. Then on compounds from ChEMBL, the public literature database, which come from different laboratories and which the model has never seen. The second is the harder test, and it is what the external numbers below describe. It is a retrospective benchmark on an outside dataset with exact training compounds removed, not an untouched independent replication: this corpus has been used elsewhere in the project, and only exact compound matches were excluded, not close analogues.
| same collection, compounds withheld | different database, never seen | |
|---|---|---|
| compounds ranked | 49,211 | 10,161 |
| pairwise comparisons | 46,330,109 | 1,926,222 |
| agreement with the measured order | 0.79 | 0.50 |
| pairs called right when potencies differ by more than tenfold | 90% | 78% |
| enrichment in the top tenth of the list | 5.5x | 3.6x |
Performance is not uniform and the aggregate hides that. These are the same 30 targets, split by how well the model ordered compounds it had never seen. A target's tier is the first thing to check before acting on a prediction for it.
Order these with confidence. Agreement of 0.6 or better, and the top of the list is strongly enriched in genuinely potent compounds.
| target | compounds | agreement with measured order | decisive pairs correct | enrichment in top tenth |
|---|---|---|---|---|
| AURKA | 221 | 0.87 | 94% | 3.7x |
| CDK1 | 137 | 0.81 | 92% | 7.7x |
| MTOR | 169 | 0.75 | 88% | 5.8x |
| SRC | 370 | 0.72 | 88% | 5.7x |
| PIM1 | 162 | 0.71 | 84% | 4.4x |
| ABL1 | 496 | 0.69 | 83% | 4.0x |
| CDK2 | 293 | 0.67 | 85% | 4.2x |
| BRAF | 217 | 0.63 | 84% | 4.0x |
| ERBB2 | 394 | 0.61 | 81% | 3.1x |
Clearly better than chance and worth using to prioritise, but the ordering is loose enough that the top of the list should be confirmed rather than trusted.
| target | compounds | agreement with measured order | decisive pairs correct | enrichment in top tenth |
|---|---|---|---|---|
| FGFR1 | 477 | 0.60 | 82% | 3.3x |
| MET | 467 | 0.55 | 81% | 5.5x |
| SYK | 363 | 0.53 | 83% | 6.2x |
| FLT1 | 132 | 0.51 | 78% | 3.9x |
| MAPK14 | 145 | 0.46 | 74% | 3.7x |
| PIK3CG | 454 | 0.42 | 71% | 3.6x |
| CDK9 | 442 | 0.42 | 74% | 4.3x |
| TYK2 | 487 | 0.41 | 71% | 2.8x |
| JAK3 | 377 | 0.41 | 70% | 5.0x |
Either the ordering is barely better than chance, or the top of the list is not enriched enough to act on, a target whose top tenth holds no more potent compounds than a random draw is not usable for screening however well the rest of the list is sorted. They are named rather than folded into an average.
| target | compounds | agreement with measured order | decisive pairs correct | enrichment in top tenth |
|---|---|---|---|---|
| AKT1 | 371 | 0.77 | 86% | 1.6x |
| PIK3CA | 371 | 0.58 | 78% | 1.6x |
| KDR | 274 | 0.49 | 75% | 1.1x |
| CDK4 | 214 | 0.48 | 81% | 1.0x |
| BTK | 398 | 0.40 | 71% | 2.7x |
| FLT3 | 467 | 0.37 | 67% | 3.4x |
| JAK2 | 258 | 0.34 | 66% | 4.2x |
| EGFR | 305 | 0.32 | 66% | 2.4x |
| MAPK1 | 499 | 0.31 | 66% | 1.0x |
| GSK3B | 332 | 0.30 | 63% | 1.2x |
| PDGFRB | 393 | 0.27 | 73% | 0.0x |
| PIK3CD | 476 | 0.09 | 57% | 3.3x |
The headline describes no individual target. Across 332 kinases with enough held-out pairs to measure, 260 reach 70% pairwise accuracy or better, 58 sit between 60% and 70%, and 14 fall below 60% against a 50% coin flip. Use it where it is strong; the rows below say where that is.
Arm: c51_rf_full_data. Sorted by the first metric column, best first, pairwise accuracy in ranker mode, rank correlation otherwise. Test ligands is the number of held-out compounds that target was scored on. Read it before the accuracy: a target scored on 50 pairs carries only about a dozen distinct ligands, and a high number there is not the same evidence as the same number over thousands of pairs.
| kinase | pairwise accuracy | MCC | test pairs | test ligands |
|---|---|---|---|---|
| ULK2 | 98.1% | 0.962 | 52 | 11 |
| GRK7 | 93.3% | 0.870 | 90 | 14 |
| STK25 | 93.2% | 0.864 | 73 | 13 |
| EPHB1 | 92.3% | 0.846 | 65 | 12 |
| CAMK2D | 90.4% | 0.808 | 19730 | 279 |
| MAP3K9 | 88.0% | 0.760 | 274 | 24 |
| DCLK1 | 87.4% | 0.743 | 166 | 19 |
| EPHB2 | 87.2% | 0.746 | 273 | 24 |
| CAMK2A | 86.9% | 0.739 | 1302 | 52 |
| OXSR1 | 86.9% | 0.738 | 84 | 15 |
| TSSK1B | 86.8% | 0.737 | 393 | 29 |
| RIPK1 | 86.2% | 0.724 | 19958 | 915 |
| PIM1 | 86.1% | 0.723 | 19939 | 1689 |
| PKMYT1 | 86.0% | 0.721 | 136 | 17 |
| PTK2 | 86.0% | 0.719 | 19620 | 745 |
| IGF1R | 85.8% | 0.716 | 19781 | 694 |
| CSNK1D | 85.8% | 0.716 | 19915 | 433 |
| PKN1 | 85.6% | 0.707 | 187 | 20 |
| TTK | 85.3% | 0.706 | 19911 | 741 |
| RAF1 | 85.3% | 0.706 | 19880 | 834 |
| IKBKE | 85.1% | 0.703 | 19914 | 256 |
| ROCK1 | 85.0% | 0.700 | 19952 | 1002 |
| FGFR1 | 85.0% | 0.699 | 19962 | 1176 |
| STK33 | 84.6% | 0.694 | 117 | 16 |
| CAMK1D | 84.6% | 0.692 | 689 | 38 |
| PRKCB | 84.6% | 0.692 | 19966 | 215 |
| MTOR | 84.5% | 0.689 | 19975 | 2222 |
| LTK | 84.4% | 0.687 | 1851 | 62 |
| CDK4 | 84.3% | 0.687 | 19963 | 1374 |
| ABL2 | 84.3% | 0.686 | 560 | 34 |
| MAPKAPK2 | 84.3% | 0.685 | 19951 | 413 |
| FGFR2 | 84.1% | 0.683 | 19917 | 777 |
| ABL1 | 84.0% | 0.680 | 19959 | 1145 |
| MST1R | 83.9% | 0.679 | 1806 | 61 |
| TBK1 | 83.8% | 0.676 | 19883 | 329 |
| MAPK9 | 83.7% | 0.674 | 19906 | 402 |
| MAPKAPK3 | 83.6% | 0.673 | 165 | 19 |
| ACVRL1 | 83.6% | 0.672 | 701 | 38 |
| MAP3K10 | 83.6% | 0.672 | 268 | 24 |
| PRKCE | 83.4% | 0.669 | 6420 | 114 |
| PIK3CA | 83.4% | 0.668 | 19970 | 2066 |
| WEE1 | 83.4% | 0.667 | 19905 | 301 |
| AXL | 83.4% | 0.667 | 19889 | 468 |
| MAP4K1 | 83.3% | 0.667 | 19326 | 247 |
| RIOK1 | 83.3% | 0.666 | 78 | 13 |
| FGFR3 | 83.3% | 0.666 | 19952 | 910 |
| PAK1 | 83.3% | 0.666 | 10665 | 147 |
| PIK3R4 | 83.3% | 0.665 | 19923 | 375 |
| CDK2 | 83.2% | 0.665 | 19901 | 2211 |
| ALK | 83.2% | 0.664 | 19930 | 697 |
| GSK3A | 83.0% | 0.661 | 19956 | 772 |
| CDK3 | 83.0% | 0.660 | 666 | 37 |
| HIPK1 | 83.0% | 0.659 | 376 | 28 |
| FGR | 82.8% | 0.656 | 3641 | 86 |
| JAK2 | 82.8% | 0.656 | 19976 | 3500 |
| MARK1 | 82.8% | 0.651 | 87 | 14 |
| PRKD3 | 82.7% | 0.655 | 10463 | 147 |
| TXK | 82.6% | 0.653 | 662 | 37 |
| ROS1 | 82.5% | 0.651 | 4922 | 100 |
| MAP4K2 | 82.5% | 0.650 | 8992 | 136 |
| NTRK3 | 82.5% | 0.650 | 10342 | 145 |
| BRAF | 82.4% | 0.648 | 19917 | 1455 |
| GRK5 | 82.4% | 0.647 | 574 | 35 |
| SYK | 82.3% | 0.646 | 19712 | 1497 |
| CAMK1 | 82.3% | 0.647 | 249 | 23 |
| NTRK2 | 82.3% | 0.646 | 19918 | 213 |
| PRKD1 | 82.2% | 0.645 | 348 | 27 |
| AKT3 | 82.1% | 0.642 | 19909 | 317 |
| CLK1 | 82.0% | 0.640 | 19959 | 268 |
| CSF1R | 82.0% | 0.640 | 19913 | 603 |
| PDGFRB | 82.0% | 0.639 | 19968 | 636 |
| TYK2 | 81.9% | 0.639 | 19982 | 1259 |
| EGFR | 81.9% | 0.637 | 19959 | 2441 |
| CAMK2G | 81.8% | 0.637 | 3884 | 90 |
| MELK | 81.8% | 0.636 | 19858 | 252 |
| PIK3CB | 81.8% | 0.636 | 19965 | 829 |
| PRKCA | 81.8% | 0.636 | 19955 | 256 |
| CLK4 | 81.8% | 0.636 | 19836 | 353 |
| CSNK2A1 | 81.8% | 0.635 | 19956 | 449 |
| EPHA6 | 81.8% | 0.632 | 252 | 23 |
| ERN1 | 81.7% | 0.632 | 153 | 18 |
| AKT1 | 81.7% | 0.633 | 19957 | 993 |
| ATM | 81.7% | 0.633 | 5619 | 107 |
| CDC42BPG | 81.5% | 0.626 | 65 | 12 |
| TSSK2 | 81.5% | 0.594 | 54 | 11 |
| EPHA2 | 81.4% | 0.628 | 16295 | 187 |
| MAPK8 | 81.4% | 0.628 | 19948 | 635 |
| SRC | 81.3% | 0.626 | 19960 | 1112 |
| PIK3CG | 81.3% | 0.626 | 19966 | 1394 |
| PTK6 | 81.3% | 0.625 | 934 | 44 |
| MAPK13 | 81.2% | 0.624 | 1097 | 48 |
| JAK1 | 81.2% | 0.624 | 19930 | 2951 |
| FLT4 | 81.2% | 0.624 | 19871 | 216 |
| MAP3K11 | 81.1% | 0.620 | 206 | 21 |
| PRKCQ | 81.0% | 0.620 | 19972 | 681 |
| INSR | 80.9% | 0.618 | 19804 | 301 |
| EPHA8 | 80.9% | 0.618 | 461 | 31 |
| DMPK | 80.9% | 0.614 | 68 | 13 |
| PRKCD | 80.8% | 0.616 | 19899 | 294 |
| PRKCG | 80.8% | 0.615 | 4848 | 100 |
| DAPK1 | 80.7% | 0.607 | 135 | 17 |
| PRKCH | 80.7% | 0.615 | 5020 | 101 |
| MAPK10 | 80.7% | 0.613 | 19939 | 562 |
| NTRK1 | 80.6% | 0.612 | 19966 | 982 |
| AKT2 | 80.5% | 0.610 | 19916 | 329 |
| GSK3B | 80.5% | 0.610 | 19946 | 1383 |
| RPS6KA3 | 80.4% | 0.608 | 19714 | 231 |
| KDR | 80.4% | 0.607 | 19934 | 3026 |
| CSNK1E | 80.4% | 0.607 | 19968 | 276 |
| MINK1 | 80.3% | 0.607 | 4977 | 102 |
| PRKG2 | 80.2% | 0.605 | 875 | 43 |
| MAP2K2 | 80.2% | 0.603 | 1209 | 50 |
| TAOK3 | 80.1% | 0.603 | 453 | 32 |
| MAPK12 | 80.1% | 0.602 | 4249 | 93 |
| PRKDC | 80.0% | 0.599 | 19860 | 234 |
| CSNK2A2 | 79.9% | 0.598 | 4643 | 97 |
| DYRK1A | 79.9% | 0.597 | 19930 | 660 |
| LRRK2 | 79.8% | 0.597 | 19965 | 1048 |
| MAP4K4 | 79.8% | 0.597 | 19713 | 276 |
| STK24 | 79.8% | 0.600 | 104 | 15 |
| TEC | 79.8% | 0.594 | 816 | 41 |
| TYRO3 | 79.8% | 0.595 | 19896 | 252 |
| AURKB | 79.7% | 0.593 | 19908 | 876 |
| EIF2AK4 | 79.6% | 0.596 | 54 | 11 |
| FGFR4 | 79.6% | 0.592 | 19951 | 606 |
| HCK | 79.6% | 0.591 | 12800 | 161 |
| IRAK4 | 79.5% | 0.591 | 19940 | 1020 |
| EPHA1 | 79.5% | 0.589 | 297 | 25 |
| AURKC | 79.4% | 0.588 | 592 | 35 |
| PRKAA1 | 79.4% | 0.588 | 12487 | 159 |
| PIM2 | 79.3% | 0.587 | 19964 | 1117 |
| JAK3 | 79.3% | 0.586 | 19973 | 1950 |
| TNK2 | 79.2% | 0.584 | 12050 | 156 |
| PLK2 | 79.1% | 0.583 | 1319 | 52 |
| PDK2 | 79.1% | 0.583 | 11422 | 152 |
| LIMK1 | 79.1% | 0.583 | 19850 | 202 |
| CDK5 | 79.1% | 0.582 | 19913 | 507 |
| CDK1 | 79.1% | 0.582 | 19957 | 1008 |
| MARK3 | 79.0% | 0.581 | 3463 | 85 |
| PDGFRA | 79.0% | 0.580 | 19938 | 432 |
| CDK19 | 79.0% | 0.580 | 4720 | 98 |
| SIK3 | 79.0% | 0.579 | 19949 | 219 |
| SIK1 | 78.9% | 0.579 | 19981 | 231 |
| STK4 | 78.8% | 0.578 | 170 | 19 |
| MAP3K12 | 78.8% | 0.576 | 15007 | 174 |
| HIPK3 | 78.8% | 0.575 | 231 | 22 |
| PLK4 | 78.7% | 0.574 | 19931 | 255 |
| MET | 78.6% | 0.573 | 19940 | 1211 |
| EPHB4 | 78.6% | 0.573 | 16707 | 187 |
| NEK4 | 78.6% | 0.572 | 1262 | 52 |
| MAPK1 | 78.6% | 0.572 | 19976 | 1110 |
| YES1 | 78.6% | 0.572 | 4262 | 93 |
| PTK2B | 78.6% | 0.571 | 7247 | 121 |
| MERTK | 78.6% | 0.571 | 19918 | 310 |
| PKN2 | 78.5% | 0.571 | 3178 | 81 |
| ERBB4 | 78.4% | 0.568 | 5840 | 109 |
| STK17B | 78.4% | 0.599 | 88 | 14 |
| FLT1 | 78.4% | 0.568 | 19889 | 1128 |
| BRSK1 | 78.4% | 0.568 | 1522 | 57 |
| SRMS | 78.3% | 0.566 | 978 | 45 |
| LCK | 78.3% | 0.566 | 19928 | 624 |
| ERBB2 | 78.3% | 0.565 | 19905 | 1330 |
| MKNK2 | 78.2% | 0.564 | 19919 | 399 |
| ACVR1 | 78.1% | 0.563 | 7828 | 126 |
| RIOK2 | 78.1% | 0.561 | 64 | 12 |
| CHEK1 | 78.1% | 0.562 | 19949 | 639 |
| PRKX | 78.0% | 0.561 | 3554 | 86 |
| FLT3 | 78.0% | 0.559 | 19956 | 1152 |
| AURKA | 78.0% | 0.559 | 19959 | 1247 |
| CLK2 | 78.0% | 0.559 | 19888 | 313 |
| MAPK14 | 77.6% | 0.552 | 19976 | 1862 |
| PIK3CD | 77.5% | 0.551 | 19860 | 1225 |
| MAP4K3 | 77.3% | 0.549 | 119 | 16 |
| TEK | 77.3% | 0.545 | 16436 | 182 |
| KIT | 77.2% | 0.544 | 19906 | 657 |
| PLK1 | 77.1% | 0.543 | 19942 | 341 |
| PRKCI | 77.1% | 0.542 | 9591 | 160 |
| DCLK2 | 77.0% | 0.544 | 265 | 24 |
| DYRK1B | 77.0% | 0.539 | 19911 | 271 |
| DDR2 | 76.9% | 0.538 | 1427 | 54 |
| BTK | 76.8% | 0.537 | 19961 | 1608 |
| BLK | 76.7% | 0.534 | 5580 | 107 |
| MAPK15 | 76.7% | 0.529 | 120 | 16 |
| ROCK2 | 76.4% | 0.528 | 19971 | 961 |
| MAPK3 | 76.4% | 0.530 | 250 | 23 |
| LYN | 76.4% | 0.528 | 19905 | 205 |
| PDPK1 | 76.4% | 0.528 | 19914 | 361 |
| PRKCZ | 76.4% | 0.528 | 1693 | 59 |
| ERBB3 | 76.4% | 0.527 | 55 | 11 |
| CDK6 | 76.3% | 0.526 | 19960 | 526 |
| RPS6KB1 | 76.2% | 0.523 | 19956 | 664 |
| IKBKB | 76.2% | 0.523 | 19934 | 405 |
| INSRR | 76.1% | 0.510 | 130 | 17 |
| MAP2K6 | 76.1% | 0.527 | 113 | 16 |
| PAK4 | 75.9% | 0.518 | 13603 | 166 |
| TAOK1 | 75.8% | 0.515 | 5759 | 110 |
| PDK1 | 75.7% | 0.514 | 16035 | 248 |
| DDR1 | 75.7% | 0.513 | 2478 | 71 |
| TGFBR1 | 75.7% | 0.513 | 19963 | 593 |
| ITK | 75.6% | 0.512 | 19930 | 444 |
| MKNK1 | 75.6% | 0.512 | 19950 | 371 |
| CDK13 | 75.6% | 0.510 | 90 | 14 |
| MYLK2 | 75.6% | 0.514 | 90 | 14 |
| NEK2 | 75.5% | 0.511 | 3577 | 86 |
| RET | 75.4% | 0.508 | 19982 | 770 |
| RPS6KA2 | 75.2% | 0.503 | 1704 | 59 |
| CDK7 | 75.2% | 0.503 | 19822 | 474 |
| AAK1 | 75.0% | 0.501 | 19426 | 198 |
| MAP3K14 | 75.0% | 0.499 | 19907 | 281 |
| MAP2K1 | 74.9% | 0.499 | 19955 | 340 |
| RPS6KA1 | 74.9% | 0.498 | 19920 | 416 |
| EPHB6 | 74.7% | 0.489 | 190 | 20 |
| MAP3K7 | 74.7% | 0.493 | 8545 | 132 |
| CSK | 74.6% | 0.490 | 481 | 32 |
| STK10 | 74.6% | 0.493 | 700 | 38 |
| PRKD2 | 74.5% | 0.490 | 6530 | 116 |
| PBK | 74.4% | 0.489 | 5317 | 104 |
| STK17A | 74.1% | 0.483 | 3098 | 80 |
| IRAK1 | 74.0% | 0.480 | 11915 | 156 |
| STK16 | 74.0% | 0.486 | 300 | 25 |
| CLK3 | 74.0% | 0.479 | 4269 | 93 |
| PIM3 | 73.9% | 0.478 | 19796 | 977 |
| SGK1 | 73.9% | 0.478 | 16695 | 184 |
| SIK2 | 73.8% | 0.477 | 19949 | 331 |
| PIK3C3 | 73.7% | 0.474 | 9287 | 137 |
| MARK2 | 73.7% | 0.474 | 3700 | 88 |
| TNIK | 73.3% | 0.467 | 1428 | 54 |
| FYN | 73.3% | 0.467 | 19958 | 342 |
| PI4KB | 73.3% | 0.466 | 5022 | 101 |
| HIPK2 | 73.2% | 0.464 | 4735 | 99 |
| DAPK3 | 73.2% | 0.463 | 7159 | 121 |
| CSNK1G2 | 73.1% | 0.462 | 2682 | 75 |
| CAMKK2 | 73.1% | 0.462 | 806 | 41 |
| CSNK1G3 | 73.0% | 0.460 | 837 | 42 |
| EIF2AK3 | 72.8% | 0.455 | 147 | 18 |
| DGKZ | 72.7% | 0.454 | 19891 | 270 |
| CDK9 | 72.7% | 0.453 | 19931 | 1322 |
| PRKG1 | 72.6% | 0.452 | 1599 | 58 |
| CSNK2B | 72.5% | 0.451 | 750 | 40 |
| PRKACA | 72.5% | 0.451 | 19938 | 281 |
| EPHA5 | 72.5% | 0.451 | 120 | 16 |
| CHEK2 | 72.4% | 0.448 | 13410 | 165 |
| LIMK2 | 72.3% | 0.447 | 1826 | 61 |
| SGK2 | 71.8% | 0.435 | 1283 | 52 |
| CSNK1G1 | 71.3% | 0.424 | 1354 | 53 |
| MAP2K7 | 71.1% | 0.424 | 298 | 25 |
| STK26 | 71.1% | 0.428 | 90 | 14 |
| CDK8 | 71.0% | 0.420 | 19883 | 248 |
| RIPK2 | 70.9% | 0.418 | 11239 | 151 |
| DAPK2 | 70.9% | 0.417 | 206 | 21 |
| PAK2 | 70.7% | 0.413 | 3823 | 89 |
| IRAK3 | 70.5% | 0.409 | 105 | 15 |
| CDC42BPA | 70.4% | 0.409 | 2641 | 75 |
| EPHA4 | 70.4% | 0.396 | 152 | 18 |
| NEK1 | 70.4% | 0.408 | 135 | 17 |
| CSNK1A1 | 70.3% | 0.407 | 19070 | 197 |
| MATK | 70.3% | 0.405 | 219 | 22 |
| ZAP70 | 70.3% | 0.405 | 1940 | 63 |
| IKBKG | 70.1% | 0.402 | 3790 | 88 |
| NUAK1 | 70.0% | 0.387 | 120 | 16 |
| ATR | 70.0% | 0.399 | 9021 | 135 |
| TAOK2 | 69.9% | 0.401 | 133 | 17 |
| MAPKAPK5 | 69.9% | 0.398 | 1086 | 48 |
| DYRK2 | 69.8% | 0.397 | 14159 | 169 |
| ACVR1B | 69.7% | 0.390 | 208 | 21 |
| BMPR1A | 69.7% | 0.394 | 231 | 22 |
| MAP3K13 | 69.7% | 0.345 | 66 | 12 |
| STK3 | 69.6% | 0.392 | 8684 | 133 |
| MAP3K2 | 69.2% | 0.390 | 348 | 27 |
| MAP3K20 | 69.2% | 0.385 | 2327 | 70 |
| PIP5K1C | 69.2% | 0.323 | 52 | 11 |
| BMPR2 | 69.2% | 0.395 | 91 | 14 |
| CDC7 | 69.0% | 0.381 | 19923 | 204 |
| CAMK2B | 68.5% | 0.369 | 2042 | 67 |
| DYRK3 | 68.4% | 0.369 | 8324 | 130 |
| RPS6KA5 | 68.3% | 0.366 | 1959 | 64 |
| GRK2 | 68.1% | 0.362 | 91 | 14 |
| MYLK | 67.8% | 0.355 | 450 | 31 |
| GSG2 | 67.7% | 0.351 | 325 | 26 |
| BMP2K | 67.6% | 0.348 | 349 | 27 |
| PHKG2 | 67.6% | 0.352 | 1040 | 47 |
| SLK | 67.5% | 0.351 | 4635 | 98 |
| BMPR1B | 67.4% | 0.350 | 350 | 27 |
| EPHA7 | 67.4% | 0.362 | 135 | 17 |
| DGKA | 67.4% | 0.348 | 19851 | 339 |
| MAP2K5 | 67.3% | 0.346 | 300 | 25 |
| DYRK4 | 67.0% | 0.340 | 2829 | 76 |
| PDK4 | 66.9% | 0.338 | 136 | 17 |
| NLK | 66.8% | 0.338 | 187 | 20 |
| BUB1 | 66.7% | 0.334 | 19893 | 233 |
| CDK12 | 66.6% | 0.332 | 7264 | 131 |
| MAP3K8 | 66.5% | 0.329 | 937 | 44 |
| NUAK2 | 66.4% | 0.330 | 134 | 17 |
| CDK11A | 66.3% | 0.322 | 89 | 14 |
| EIF2AK2 | 66.1% | 0.311 | 65 | 12 |
| GRK6 | 66.1% | 0.351 | 65 | 12 |
| CIT | 66.0% | 0.319 | 209 | 21 |
| RPS6KA6 | 66.0% | 0.320 | 297 | 25 |
| TNK1 | 65.9% | 0.291 | 91 | 14 |
| ULK1 | 65.9% | 0.318 | 88 | 14 |
| CDK16 | 65.7% | 0.313 | 350 | 27 |
| CHUK | 65.6% | 0.313 | 7040 | 120 |
| ULK3 | 65.6% | 0.322 | 90 | 14 |
| FRK | 65.5% | 0.311 | 19784 | 298 |
| MAP3K19 | 65.2% | 0.316 | 135 | 17 |
| EPHA3 | 65.1% | 0.268 | 66 | 12 |
| HIPK4 | 65.0% | 0.301 | 2748 | 76 |
| TIE1 | 64.9% | 0.292 | 77 | 13 |
| PLK3 | 64.9% | 0.298 | 3674 | 87 |
| MARK4 | 64.5% | 0.288 | 76 | 13 |
| MAP4K5 | 64.4% | 0.287 | 8900 | 136 |
| MAPK11 | 64.0% | 0.279 | 3223 | 81 |
| CDKL2 | 63.1% | 0.232 | 65 | 12 |
| MAP3K3 | 62.8% | 0.256 | 78 | 13 |
| FER | 62.6% | 0.253 | 1049 | 47 |
| CDK17 | 62.5% | 0.275 | 64 | 12 |
| MAP2K3 | 62.2% | 0.244 | 119 | 16 |
| FES | 61.0% | 0.217 | 523 | 33 |
| TNNI3K | 59.8% | 0.195 | 887 | 43 |
| CAMK4 | 58.8% | 0.198 | 51 | 12 |
| RPS6KA4 | 58.7% | 0.179 | 189 | 20 |
| SRPK1 | 58.1% | 0.161 | 1392 | 55 |
| GRK1 | 57.9% | 0.144 | 76 | 13 |
| GAK | 57.7% | 0.155 | 1325 | 52 |
| MAP2K4 | 57.6% | 0.100 | 66 | 12 |
| BMX | 57.3% | 0.147 | 4308 | 98 |
| TGFBR2 | 56.1% | 0.125 | 346 | 27 |
| PIP4K2C | 51.5% | 0.007 | 66 | 12 |
| MAPK7 | 50.9% | -0.014 | 55 | 11 |
| RIPK3 | 50.4% | 0.009 | 119 | 16 |
| ICK | 42.0% | -0.164 | 119 | 16 |
| CSNK2A3 | 40.7% | -0.203 | 118 | 16 |
Of the 179 targets this project classes as understudied, the published model already scores 163, because they are targets with modest data rather than targets with none. Taken at face value they look weaker: median accuracy 0.736 against 0.769 for the rest.
They also carry far less held-out data, a median of 754 pairs against 19,893, a twenty-six-fold difference. Comparing them at matched data volume, in the one band where both groups are well populated, understudied targets score 0.715 (n=31) against 0.723 (n=22): a difference of -0.0077 ± 0.0403, indistinguishable from zero.
The model answers a comparison: given one kinase sequence and two ligands, which ligand is more potent against that kinase. Ranking a compound library is a sequence of such comparisons, so this is the primitive operation of virtual screening rather than a proxy for it. Because the orientation of every pair is randomised, a constant output scores exactly 0.500 and a Matthews correlation of 0.000, so the chance line needs no argument. That is not the same as there being no baseline worth beating, and two are reported beside the model below: ranking by molecular size alone, and looking a compound's potency up from another kinase where such data exists.
Measurements are drawn from the Kinase Knowledgebase (KKB), Eidogen-Sertanty, release Q2-2026, restricted to human enzyme assays with a defined potency between pIC50 3 and 11.
From the corpus to the fitted model, every step:
| step | measurements | what changes |
|---|---|---|
| Source corpus | 841,187 | every KKB row on the 500-gene panel |
| Training half | 673,659 | 167,528 rows are held out for testing and never trained on |
| Exact values only | 417,507 | records that give only a limit, such as weaker than 10 micromolar, rather than an exact potency, are set aside; repeat measurements of the same compound against the same kinase are collapsed to their median |
| Fitted | 417,507 | all of them; nothing is held back from the fit |
The training half spans 500 kinase genes and 742 distinct sequences, including 13,218 rows carrying a mutated sequence; a mutant is represented by its own sequence rather than folded into the wild type.
Each ligand is described by a Morgan count fingerprint of radius 2 and 1,024 bits, computed with RDKit from the canonical SMILES, concatenated with fourteen physicochemical and compositional descriptors: molecular weight, heavy-atom count, bond count, rotatable bonds, ring count, and the counts of C, N, O, S, F, Cl, Br, I and P. The descriptors are part of the arm that was scored and are reported for exactness. Their measured contribution is negligible: paired per kinase over the 213 targets with at least 1,000 pairs, adding them is worth +0.0023 ± 0.0030 on count fingerprints. An earlier version of this document explained them by saying a circular fingerprint represents molecular size poorly; that explanation predicted a material gain on BINARY fingerprints, and the measured gain there was +0.0032 ± 0.0031, no different. The explanation is withdrawn and plain Morgan counts would serve as well.
Each kinase sequence is embedded with ESM2 (esm2_t12_35M_UR50D), giving one vector per residue, and reduced to a single fixed 480-dimensional vector by mean pooling over residues. Pooling is fixed rather than learned. Concatenated with the 1,038 ligand features, the vector presented to the model is 1,518 dimensions: 480 protein and 1,038 ligand.
The model is trained on human kinase sequences and ligand structures from the Kinase Knowledgebase, and on nothing else. Its training table carries 673,659 rows over 500 kinases and 742 distinct sequences, of which 13,218 rows are point-mutant sequence records. Every row is an original measurement.
The structure-free framing of this project, predicting engagement from protein sequence and ligand chemistry alone, with no three-dimensional pose, using a protein language model together with a chemical language model, follows Fondrie and colleagues, Structure-free, site-resolved contrastive learning extends small-molecule discovery beyond the reach of structure-based modeling (Talus Bioscience, bioRxiv 2026), which introduces Ptarmigan-1.
The debt is specific and worth stating precisely. This project's neural arm was adapted from that architecture: frozen ESM2 residue embeddings and frozen ChemBERTa ligand embeddings projected into a shared space, with the temperature-scaled softmax pooling over residues that Ptarmigan-1 uses to turn residue-level scores into a protein-level call reused here as attention pooling, and its leakage-safe Bemis-Murcko scaffold-split discipline. Trainable projections and head only, both backbones frozen, as a lighter-weight analogue of its LoRA fine-tuning.
The model reported here is not that model. The neural arm was beaten on every measure by the random forest described above, which uses fixed mean pooling and a count fingerprint rather than learned pooling and a chemical language model, and which is what all results in this document describe. The two also answer different questions: Ptarmigan-1 predicts and localises engagement across the proteome, including at cryptic and disordered sites; this model ranks ligands by potency against a kinase for which measured data already exists.
A random forest regressor of 200 trees (minimum two samples per leaf) is fitted pointwise on (protein vector, ligand features) against measured pIC50. Two ligands are then compared by the difference of the model's two predictions, so that P(A more potent than B) follows the sign of sA − sB. This construction is exactly antisymmetric: exchanging the two ligands reverses the prediction identically, and the model cannot return contradictory answers for the same pair presented in a different order. A forest given both ligands as joint input would carry no such guarantee. It also means a library of n compounds is scored in n forward passes rather than n2 comparisons.
This model has no training pairs. The forest is fitted pointwise on individual pIC50 values; pairs exist only at evaluation, and every non-tied pair of ligands measured against the same sequence is scored. There is no gap filter on pair construction.
Results are then reported in three bands of true potency difference: under half a log, half to one log, and more than one log. The band most worth quoting is the last, because that is the separation a chemist acts on. For context on what a resolvable difference is, disagreement between two independent publications reporting the same compound against the same kinase has a 90th percentile of 0.715 log units over 204,438 comparisons, rising to 1.218 after deduplication; the source file recommends treating one log as the floor. Pairs closer than that are harder rather than impossible: the model still scores 57.8 percent under half a log, against a 50 percent coin flip.
The advantage that cross-kinase knowledge confers is not assumed away, it is measured. The compound-lookup baseline above is exactly that advantage made explicit, predicting from the same compound’s potency against another kinase, and it scores 73.1 percent. The model scores 76.8 percent on those same pairs, so it adds information beyond knowing the compound elsewhere. What these internal numbers do not measure is performance on chemistry absent from the corpus entirely; the external section speaks to that.
Ligands are withheld by Bemis-Murcko scaffold, keyed on the (sequence, ligand) pair. Every kinase sequence remains in training. No protein and no sequence family is withheld, because a kinase for which no data exists is not the deployment scenario: the kinome is among the most exhaustively characterised regions of the proteome. Evaluation covers 2,965,273 held-out ligand pairs across 413 targets. The kinase is the unit of replication for the kinase-level mean and its standard error, and only for those. Pooled accuracy, the margin deciles and the selectivity figures are computed over pairs that share compounds and are therefore not independent observations; treat them as descriptive rather than as quantities with an interval.
The headline results are obtained without three bodies of work that are built and available: 4,555 CANDIDATE orthologue sequences across 475 organisms, screened by binding-site identity; per-residue contact maps from 1,417 solved co-complex structures covering 170 targets, with 977 identified ligands and 182 sites excluded because a nucleotide or crystallisation additive rather than an inhibitor defined them; and 821,370 BindingDB measurements retained after removing every compound-target pair appearing in the frozen ChEMBL benchmarks. These are headroom, not caveats.
Kinase Knowledgebase (KKB), Eidogen-Sertanty, release Q2-2026, the training and evaluation corpus: eidogen-sertanty.com/kinasekbmarvin.php.
BindingDB, ingested and held for future arms, not used in the results reported here. Gilson et al., Nucleic Acids Res. 44:D1045 (2016), doi:10.1093/nar/gkv1072; bindingdb.org.
ChEMBL, used only as a frozen external benchmark, never trained on; every compound-target pair appearing in it was removed from the BindingDB import. Zdrazil et al., Nucleic Acids Res. 52:D1180 (2024), doi:10.1093/nar/gkad1004; ebi.ac.uk/chembl.
RCSB Protein Data Bank, co-complex structures used for the binding-site analysis reported separately. Berman et al., Nucleic Acids Res. 28:235 (2000), doi:10.1093/nar/28.1.235; rcsb.org.
UniProt, sequence and accession mapping. UniProt Consortium, Nucleic Acids Res. 51:D523 (2023), doi:10.1093/nar/gkac1052; uniprot.org.
ESM2 protein language model, checkpoint facebook/esm2_t12_35M_UR50D, 35M parameters, 12 layers, 480-dimensional per-residue representations. Lin et al., Science 379:1123 (2023), doi:10.1126/science.ade2574; code at github.com/facebookresearch/esm.
ChemBERTa chemical language model, checkpoint seyonec/ChemBERTa-zinc-base-v1, 768-dimensional embeddings, used in the comparison arm. Chithrananda et al., arXiv:2010.09885 (2020), arxiv.org/abs/2010.09885.
Morgan / ECFP count fingerprints, radius 2, 1,024 bits, and all physicochemical descriptors computed with RDKit: rdkit.org. Method: Rogers & Hahn, J. Chem. Inf. Model. 50:742 (2010), doi:10.1021/ci100050t.
Bemis-Murcko scaffolds, used to define the held-out ligand split. Bemis & Murcko, J. Med. Chem. 39:2887 (1996), doi:10.1021/jm9602928.
Random forest regressor, scikit-learn: scikit-learn.org. Breiman, Machine Learning 45:5 (2001), doi:10.1023/A:1010933404324.
Pairwise ranking formulation: Burges et al., Learning to Rank using Gradient Descent, ICML 2005, doi:10.1145/1102351.1102363.
Matthews correlation coefficient: Matthews, Biochim. Biophys. Acta 405:442 (1975), doi:10.1016/0005-2795(75)90109-9.
A model cannot be measured on data it was trained on. Maximum evidence and maximum data therefore cannot be the same object, so both are released rather than quietly choosing one.
kfm_ranker_v1_final, was trained on 417,507 measurements. The FRONTIER model, bundle kfm_ranker_production, was trained on 780,067 measurements, which is every measurement in the corpus. If the model you are holding was trained on roughly four hundred thousand measurements it is the VALIDATED one; roughly eight hundred thousand means it is the FRONTIER one. Each bundle states its own count in config.json under trained_on, and its own release name under release. Check that field rather than the directory name.| VALIDATED kfm_ranker_v1_final | FRONTIER kfm_ranker_production | |
|---|---|---|
| measurements trained on | 417,507 | every measurement held |
| data used | training half, exact values only | both halves, exact values and limit-only records |
| architecture | identical: 200 trees, same features, same seed, same training method | |
| measured accuracy | 74.0% per target, 90.3% on decisive pairs | none, and none is possible |
| use it when | you need a number you can defend | you want the most informed prediction available |
Point it at a kinase you already have data on and a compound library you have not measured. It scores every compound in one pass and returns a ranked list, each comparison carrying a confidence. Act on the confident fraction: accuracy rises steeply as you narrow to the calls the model is most certain about, so the operating point is yours to choose.
Check the target first. This is the single most important step and it is not optional. Performance varies more between targets than it does between any two models we have tried. On its strongest targets AURKA, CDK1, MTOR, SRC, PIM1, ABL1, CDK2, BRAF, it calls decisive pairs correctly around nine times in ten on chemistry from a source it has never seen. On its weakest it is close enough to a coin flip that a prediction should carry no weight. The tier table above is the guide, and there is no aggregate number that substitutes for reading it.
It ranks; it does not predict absolute potency. Two compounds whose true potencies differ by less than two publications routinely disagree by cannot be reliably adjudicated from this assay corpus, and the model is correctly close to chance on them rather than confidently wrong.
Where a compound already has measured potency against a related kinase, use that measurement: it is direct evidence. It is not, however, a stronger predictor than the model on this benchmark, that lookup scores 73.1 percent where it is defined, against the model's 76.8 percent on the same pairs. The lookup was undefined on 49.7 percent of evaluation pairs, and it is there that the model is the only option.
It is a triage tool that runs before docking or crystallography, not a replacement for either, and not a substitute for measurement.
It sorts a compound library against a kinase you already have data on, and reports a margin on each call that tracks how often it is right. Used as a filter on its confident predictions it is right 88 percent of the time on chemistry from a source it has never seen. A coin flip is 50%.
Scores a pair in milliseconds with no structure needed, against minutes to hours for docking, so it triages a library before any structure-based work begins. Structural evidence behind the binding-site work.
The complete experiment record, including every claim that did not
survive review, is kept separately in collab/cycles/ and is deliberately not
part of this document.
Generated from collab/cycles/*/claude_manifest.json by scripts/37_live_report.py · 08 August 2026 at 23:05