KinaseFoundationModel
Model 1 · potency ranking

Rank compounds against one kinase

ligand A sequence ligand B

Pick a kinase, supply two or more compounds, and the model orders them for that kinase. Unlike version 1, nothing is scored on its own and then subtracted: the forest is handed the whole comparison as one row: ligand A, the sequence, ligand B, in that precise order, and returns the probability that ligand A is the more potent of the two. Every pair of the compounds you supply is put to it that way, and the ranking below is a summary of those answers.

The potency model: ligand A, one kinase sequence and ligand B enter a single random forest in that fixed order. Each ligand becomes a 1,024-bit Morgan count fingerprint plus 14 descriptors; the sequence becomes 480 ESM2 numbers. The forest returns the probability that ligand A is the more potent of the two, shown on a confidence bar running from A binds tighter to B binds tighter. The worked case is bosutinib, measured pIC50 8.96, against a pyrazolo[3,4-d]pyrimidine at 4.50 on the ABL1 kinase domain, RCSB 3UE4.
The whole comparison is one row. The order is the question: ligand A first, the sequence in the middle, ligand B last. Every pair you supply below goes through exactly this, in both ligand orders, and the two answers are averaged. Figure taken from the LigASeqLigB report.

1 · Kinase target

One kinase at a time. none selected
Paste the full-length UniProt sequence. A sequence the model was not trained on is scored and flagged : what that does and does not buy you.

2 · Compounds

Two at minimum, 5 at most per run, drawn and pasted combined. Every pair is scored, so n compounds is n(n − 1)/2 comparisons, 10 compounds is 45.

You do not have to draw it. To paste a SMILES string, use the editor's Open Structure button (the folder icon, top left) and choose Paste from clipboard, it accepts SMILES and will draw the molecule for you. Then press the button below to read it back out.
You do not have to draw it. To paste a SMILES string, use the editor's Open Structure button (the folder icon, top left) and choose Paste from clipboard, it accepts SMILES and will draw the molecule for you. Then press the button below to read it back out.
We send the full results there, and let you know when the models change.
 up to 5 structures per run

3 · How to read the confidence

This model's prediction strength is the larger of the two output probabilities, so it runs from 0.5, a coin flip, to 1.0. These are the accuracies measured for each band on the ChEMBL test set, and they are what the colours below mean. The selectivity tool defines strength the same way, but its accuracies were measured on a different test set, so the same strength does not buy the same accuracy on both.

0.90 – 1.00 right 87.2% of the time, 3.0% of comparisons
0.80 – 0.90 right 85.0% of the time, 5.3%
0.70 – 0.80 right 83.2% of the time, 12.2%
0.60 – 0.70 right 74.6% of the time, 26.9%
0.50 – 0.60 right 58.4% of the time, 52.7%

What these bands mean in practice, and where the model is weakest: potency limitations.

The cutoff does not change the order, every comparison counts towards the score. It reports how many comparisons were too close to call, which is the honest measure of how firm the ranking is: the more of them, the more the neighbouring rows should be read as roughly equivalent.

4 · Run

Method, training set, test set and every figure quoted here: LigASeqLigB v2 potency report.