KinaseFoundationModel

Limitations of version 1

Version 1 is the original pointwise scorer: one kinase and one ligand per row, returning a number on the pIC50 scale, with comparisons made afterwards by subtracting two independent scores. It remains published in full. This page states what it does not do.

Looking for the current release? The two version 2 comparator models have their own limitations page, and their confidence scales are not interchangeable with version 1's margin or with each other.

Stated plainly

Quote it per target, never in aggregate. External per-target rank correlation runs from 0.09 to 0.83. A minority of targets perform at or below chance. An average across the panel hides that spread, so the tool reports each selected kinase separately rather than pooling them.
Absolute potency does not transfer; ordering does. Only 3.9% of the training data is inactive, because the literature reports exact numbers for compounds that bind and vague ones for compounds that do not. The model therefore predicts roughly two log units too potent on chemistry outside its corpus. Use it to rank, not to estimate an IC50.
It does not yet generalise to kinases it has never seen. Ranking compounds for a kinase outside the 478-member panel measured at 0.48 to 0.52, indistinguishable from chance. You may paste any sequence into the tool, and it will be scored, but anything outside the panel is labelled as unvalidated and should be treated that way.
It cannot answer selectivity. Each target's scores sit on that target's own scale, so subtracting two of them reports the difference between two scales as if it were a preference between two kinases. That question is what the version 2 selectivity model was built for; version 1 should not be used for it.

Two further results worth knowing. On compound pairs whose ranking flips between two kinases, the selectivity question, ligand information alone scores exactly 0.500 while the full model reaches 0.712, which is the clearest evidence that the protein half is carrying real signal. Against that, a simple baseline that looks up a compound’s average potency in the training set scores 0.731 on the shared benchmark versus the model’s 0.723, so on the headline metric the model does not yet beat remembering which compounds are generally potent. Both numbers are in the full results.

Where it works, target by target

Performance is not uniform, so an average across the panel would be misleading. These are held-out ChEMBL compounds, structures the model never saw, scored per target. Agreement is the rank correlation with the measured order; decisive pairs are those whose measured potencies differ by more than tenfold.

Reliable

Order these with confidence. Agreement of 0.6 or better, and the top of the list is strongly enriched in genuinely potent compounds.

TargetCompoundsAgreement with measured orderDecisive pairs correctEnrichment in top tenth
AURKA2190.8895%2.7×
MTOR1690.7789%4.7×
ABL14960.7385%3.2×
SRC3700.7388%5.4×
CDK22910.7187%5.5×
CDK11370.7186%7.7×
BRAF2160.6485%4.5×
FGFR14770.6484%3.7×
PIM11610.6480%4.4×
ERBB23920.6180%3.1×

Useful for triage

Clearly better than chance and worth using to prioritise, but the ordering is loose enough that the top of the list should be confirmed rather than trusted.

TargetCompoundsAgreement with measured orderDecisive pairs correctEnrichment in top tenth
MET4670.5782%5.1×
JAK33770.5277%5.5×
SYK3630.5183%5.6×
CDK94410.4978%4.1×
FLT11320.4776%3.9×
MAPK141380.4675%3.5×
PIK3CG4490.4372%3.8×
TYK24870.4271%4.7×

Do not rely on these

Either the ordering is barely better than chance, or the top of the list is not enriched enough to act on. A target whose top tenth holds no more potent compounds than a random draw is not usable for screening however well the rest of the list is sorted. They are named rather than folded into an average.

TargetCompoundsAgreement with measured orderDecisive pairs correctEnrichment in top tenth
AKT13710.7585%1.1×
PIK3CA3640.5576%1.7×
CDK42130.4881%1.0×
KDR2720.4875%1.1×
FLT34650.3968%4.2×
BTK3970.3970%2.2×
PDGFRB3930.3676%0.0×
MAPK14980.3367%1.0×
JAK22550.3165%4.1×
GSK3B3180.3165%3.1×
EGFR3050.2663%2.7×
PIK3CD4750.1459%2.7×

Full per-target results for all 437 evaluated kinases, including those that perform at or below chance, are in the model results report.

How version 1 differs from version 2

Version 1Version 2 potencyVersion 2 selectivity
What one row holdsone kinase, one ligandligand, sequence, ligandsequence, ligand, sequence
What it returnsa score on the target's own scaleprobability that ligand A is more potentprobability that protein A binds more potently
How two things are compareda wrapper subtracts two independent scoresthe forest is trained on the comparisonthe forest is trained on the comparison
Answers selectivity?no: scales are not comparable across targetsno: one target at a timeyes, this is what it is for
Confidence signalmargin, a monotone transform, not calibratedprediction strength, 0.5 to 1.0prediction strength, 0.5 to 1.0

Version 2 is a different architecture, not a retrain, and its numbers are not comparable to version 1's line by line. See the version 2 overview.