Choose a kinase on the left and a kinase on the right, draw or import the compounds
that go between them, and the model says which of the two proteins binds each compound
more potently. The two proteins occupy fixed positions and the position is the
question: the forest is handed sequence A, the ligand, then sequence B as a single row
and returns two probabilities that sum to 1.
This is the question version 1 could not answer.
The version 1 scorer puts each target on its own scale, so the gap between two of its
scores reports scale differences as if they were selectivity. This model is trained
on the comparison itself, so the answer does not depend on the two targets sharing a
scale.
The two proteins occupy fixed positions and the position is the question. Figure
taken from the
SeqALigSeqB report;
structures are ABL1 from PDB 2GQG, which has dasatinib bound, and GSK3B from PDB 1Q5K.
1 · The kinases, and the compounds to put between them
Pick two to five kinases. Every pair among them is scored against every compound
you supply, and the result is a grid showing which kinase each compound prefers.
KinasesChoose two to five
none selected ·
Or add your own protein sequence
LigandDraw or paste a compound
You do not have to draw it. To paste a SMILES string, use the editor's Open Structure button (the folder icon, top left) and choose Paste from clipboard, it accepts SMILES and will draw the molecule for you. Then press the button below to read it back out.
Or add more below: drawn and pasted combined, up to 5 per run.
Pasting works, including for mutants, and overrides the kinase chosen on that side.
Use the full-length UniProt sequence, not a kinase-domain slice :
what a pasted sequence changes.
2 · More compounds (optional)
We send the full results there, and let you know when the models change.
up to 5 structures per run, drawn and pasted combined
One compound is enough for this model, the comparison is between the two
kinases, not between the compounds. A list simply asks the same question of each
compound in turn.
3 · How to read the confidence
This model's prediction strength is the larger of the two output probabilities,
so it runs from 0.5, a coin flip, to 1.0. That is the same definition the potency tool
uses, so the two can be read side by side, but what a given strength buys is
specific to each model, so use the numbers below rather than carrying a threshold
across. These are the accuracies measured for each band on the ChEMBL test set.
0.9 – 1.0 right 99.3% of the time, 4.4% of comparisons 0.8 – 0.9 right 96.4% of the time, 11.8% 0.7 – 0.8 right 88.4% of the time, 20.7% 0.6 – 0.7 right 74.4% of the time, 28.9% 0.5 – 0.6 right 57.9% of the time, 34.3%
0.70 is the operating point the report argues for.
Acting only at 0.70 and above answers more than a third of comparisons at 92.3%,
against 75.3% if you answer everything. Going on to 0.90 buys only 7 more points and
costs 88% of what remained. For a shortlist that must not contain mistakes, 0.80
answers 505,499 comparisons at 97.2%.
A third of comparisons land in the bottom band.
Two kinases that a compound genuinely cannot tell apart are two kinases the model
cannot tell apart either: where the true separation is under half a log, test accuracy
is 57.7%. Strength also does not mean the same thing on the two sets, a 0.70 is
right 99% of the time on training comparisons and 92% on ChEMBL, so the cutoff
has to be read off the test column above.
Anything weaker than the cutoff is reported as too close to call rather than
given a winner. Nothing is hidden: the counts and the two probabilities are still
shown for every compound, and the number withheld is stated with the results.