Where the training data comes from, and what sequence each measurement is attached to

Named publications, real sequences, and the exact training rows they produced. Written to be checked, not believed.

Generated 2026-08-08 22:45 from data/provenance_audit.json, which queries KKB directly. Every count below is read from that file.

Why this report exists

A kinase inhibitor that is potent against a wild-type kinase can be useless against a point mutant of the same kinase, and that difference is often the entire subject of the paper reporting it. So if a mutant measurement is filed against the wild-type sequence, the model is shown two contradictory potencies for one protein and the correct thing for it to learn is to ignore the protein.

Why the evidence below is sequences and identifiersA claim that measurements are filed against the right protein is only worth as much as the evidence behind it. So this page gives named publications, the exact sequence each measurement was attached to, and the substitution shown in place, rather than a summary statistic asking to be believed.

The publications below were selected by a search over the mutation-rich kinases and are then pinned, because re-running that search across the full panel is prohibitively slow. Each one is verified end to end here, so the selection affects which examples you see and not whether they hold.

The worked examples

Example 1. publication reporting BOTH wild type and mutant
Tilting the Scales toward EGFR Mutant Selectivity: Expanding the Scope of Bivalent Type V Kinase Inhibitors.
J Med Chem 2024 · target EGFR · wild-type sequence 1,210 residues, md5 99d03b567dbc

L858R: the sequence actually attached to these 36 measurements

residues 834 to 882, substitution L858R
wild typeVHRDLAARNVLVKTPQHVKITDFGLAKLLGAEEKEYHAEGGKVPIKWMA
mutantVHRDLAARNVLVKTPQHVKITDFGRAKLLGAEEKEYHAEGGKVPIKWMA

L858R;T790M: the sequence actually attached to these 35 measurements

residues 766 to 814, substitution T790M
wild typeMASVDNPHVCRLLGICLTSTVQLITQLMPFGCLLDYVREHKDNIGSQYL
mutantMASVDNPHVCRLLGICLTSTVQLIMQLMPFGCLLDYVREHKDNIGSQYL
residues 834 to 882, substitution L858R
wild typeVHRDLAARNVLVKTPQHVKITDFGLAKLLGAEEKEYHAEGGKVPIKWMA
mutantVHRDLAARNVLVKTPQHVKITDFGRAKLLGAEEKEYHAEGGKVPIKWMA

L858R;T790M;C797S: the sequence actually attached to these 23 measurements

residues 766 to 814, substitution T790M
wild typeMASVDNPHVCRLLGICLTSTVQLITQLMPFGCLLDYVREHKDNIGSQYL
mutantMASVDNPHVCRLLGICLTSTVQLIMQLMPFGCLLDYVREHKDNIGSQYL
residues 773 to 821, substitution C797S
wild typeHVCRLLGICLTSTVQLITQLMPFGCLLDYVREHKDNIGSQYLLNWCVQI
mutantHVCRLLGICLTSTVQLITQLMPFGSLLDYVREHKDNIGSQYLLNWCVQI
residues 834 to 882, substitution L858R
wild typeVHRDLAARNVLVKTPQHVKITDFGLAKLLGAEEKEYHAEGGKVPIKWMA
mutantVHRDLAARNVLVKTPQHVKITDFGRAKLLGAEEKEYHAEGGKVPIKWMA

What this publication contributed

mutation as recorded in KKBresolved substitutionrows in this papercompoundstraining rows for this formsequence md5
L858RL858R36321,113f02e295764e7
L858R;T790MT790M, L858R353154119ebead2a842
L858R;T790M;C797ST790M, C797S, L858R232358178442f2fae03
WTwild type343013,12199d03b567dbc
Label corrected by the tracePinned as mutant_only by a protocol-level search; the row-level trace finds 3 mutant and 1 wild-type groups, so it is reported as both.
What this provesThis publication contributed 4 distinct sequences for one gene, and 128 of its rows reach training. The wild-type rows carry the wild-type sequence and each mutant carries its own, and deduplication did not collapse them into one another. If the historical bug were present, this number would be 1 and every row would show the wild-type md5.
Example 2. publication reporting ONLY mutant data
Identification of M4205A Highly Selective Inhibitor of KIT Mutations for Treatment of Unresectable Metastatic or Recurrent Gastrointestinal Stromal Tumors.
J Med Chem 2023 · target KIT · wild-type sequence 976 residues, md5 f753f2b2d975

V654A: the sequence actually attached to these 26 measurements

residues 630 to 678, substitution V654A
wild typeHLTEREALMSELKVLSYLGNHMNIVNLLGACTIGGPTLVITEYCCYGDL
mutantHLTEREALMSELKVLSYLGNHMNIANLLGACTIGGPTLVITEYCCYGDL

What this publication contributed

mutation as recorded in KKBresolved substitutionrows in this papercompoundstraining rows for this formsequence md5
V654AV654A2626304c7df54462e46
exon 11/13;(544-976,V559D;V654A)unparseable660
exon 11/14;(544-976,V559D;T670I)unparseable660
exon 11/17;(544-976,V560G;D816V)unparseable660
exon 11/17;(544-976,V560G;N822K)unparseable660
exon 11;(544-976);wildtypeunparseable660
exon 11;(544-976,557-558)unparseable660
exon 11;(544-976,V559A)unparseable660
exon 11;(544-976,V559D)unparseable660
exon 11;(544-976,V560G)unparseable660
exon 13;(544-976;K642E)unparseable660
exon 13;(544-976;V654A)unparseable660
exon 14;(544-976;T670I)unparseable660
exon 17;(544-976;A829P)unparseable660
exon 17;(544-976;D816E)unparseable660
exon 17;(544-976;D816F)unparseable660
exon 17;(544-976;D816H)unparseable660
exon 17;(544-976;D816I)unparseable660
exon 17;(544-976;D816V)unparseable660
exon 17;(544-976;D816Y)unparseable660
exon 17;(544-976;D820E)unparseable660
exon 17;(544-976;D820Y)unparseable660
exon 17;(544-976;Y823D)unparseable660
132 of this publication's rows were dropped, not mislabelledThey name variants that change the length of the protein, such as exon-19 deletions or internal tandem duplications, which cannot be expressed as a substitution on a fixed-length sequence. They are discarded rather than filed against the wild type. That is the safe failure, but it is still a loss and it is counted here rather than left silent.
What this provesEvery row this publication contributed carries a sequence whose md5 differs from the wild type, at exactly the substituted positions highlighted above. If the historical bug were still present, all of these rows would show the wild-type md5 and the model would be trained on two contradictory examples of the same protein.
Example 3. publication whose variant data is unrepresentable, so only its wild-type rows survive
Discovery of potent and selective HER2 inhibitors with efficacy against HER2 exon 20 insertion-driven tumors, which preserve wild-type EGFR signaling.
Nat Cancer 2022 · target EGFR · wild-type sequence 1,210 residues, md5 99d03b567dbc

What this publication contributed

mutation as recorded in KKBresolved substitutionrows in this papercompoundstraining rows for this formsequence md5
(empty, wild type)wild type2213,12199d03b567dbc
WTwild type484713,12199d03b567dbc
del19unparseable48470
del19,T790Munparseable48470
Label corrected by the tracePinned as both by a protocol-level search; the row-level trace finds 0 mutant and 2 wild-type groups, so it is reported as wild_type_only.
96 of this publication's rows were dropped, not mislabelledThey name variants that change the length of the protein, such as exon-19 deletions or internal tandem duplications, which cannot be expressed as a substitution on a fixed-length sequence. They are discarded rather than filed against the wild type. That is the safe failure, but it is still a loss and it is counted here rather than left silent.
What this showsEvery variant this publication reports changes the length of the protein, so none of them can be represented. Its wild-type rows are trained on and its variant rows are dropped. Nothing is mislabelled, and nothing about the variants is learned.

The sharpest form: one compound, both proteins

Compound Cc1cc(Nc2ncnc3cnc(nc23)N4CCN(CC4)C(=O)C=C)ccc1Oc5ccc6c(c5)ncn6C was measured against both forms. Its training rows:

formpIC50classsequence md5
WT5.54Low99d03b567dbc
The same compound appears against both forms, on separate rows, with different sequences. Deduplication did not collapse them.1 distinct sequences across these rows.
Example 4. publication whose variant data is unrepresentable, so only its wild-type rows survive
Cellular Context Influences Kinase Inhibitor Selectivity.
J Med Chem 2026 · target ABL1 · wild-type sequence 1,130 residues, md5 d24f1ea01ac4

What this publication contributed

mutation as recorded in KKBresolved substitutionrows in this papercompoundstraining rows for this formsequence md5
(empty, wild type)wild type336,055d24f1ea01ac4
E255K-phosphorylatedunparseable110
F317I-nonphosphorylatedunparseable110
F317I-phosphorylatedunparseable110
F317L-nonphosphorylatedunparseable110
F317L-phosphorylatedunparseable110
H396P-nonphosphorylatedunparseable110
H396P-phosphorylatedunparseable110
M351T-phosphorylatedunparseable110
Q252H-nonphosphorylatedunparseable110
Q252H-phosphorylatedunparseable110
T315I-nonphosphorylatedunparseable110
T315I-phosphorylatedunparseable110
Y253F-phosphorylatedunparseable110
nonphosphorylatedunparseable110
phosphorylatedunparseable110
Label corrected by the tracePinned as both by a protocol-level search; the row-level trace finds 0 mutant and 1 wild-type groups, so it is reported as wild_type_only.
15 of this publication's rows were dropped, not mislabelledThey name variants that change the length of the protein, such as exon-19 deletions or internal tandem duplications, which cannot be expressed as a substitution on a fixed-length sequence. They are discarded rather than filed against the wild type. That is the safe failure, but it is still a loss and it is counted here rather than left silent.
What this showsEvery variant this publication reports changes the length of the protein, so none of them can be represented. Its wild-type rows are trained on and its variant rows are dropped. Nothing is mislabelled, and nothing about the variants is learned.

Every mutant row that does not reach training, and why

The examples above show individual publications. This is the whole corpus. A record naming a mutation we cannot resolve is dropped rather than filed against the wild-type sequence, so the failure mode is lost data and never wrong data. That is the right trade, but it is still a loss, so here is its size and its composition.

categoryrowswhy it cannot be represented
insertion or deletion2,383changes the length of the protein, so it cannot be applied to a fixed-length sequence as a substitution
other unparsed1,054does not match any substitution grammar we recognise
construct or domain description640names which part of the protein was expressed, not a variant
phosphorylation state495describes the enzyme preparation, not a sequence change
internal tandem duplication264adds residues, so the sequence length changes
total dropped4,836

For scale, the training table holds 13,189 mutant rows across 237 distinct mutations, so 4,836 rows are lost against 13,189 retained.

The honest limitationMost of the loss is not a parser weakness that better code would fix. An EGFR exon-19 deletion removes residues; an internal tandem duplication adds them. This model represents a protein as a fixed-length amino-acid string with substitutions applied in place, so a variant that changes the length has nowhere to go. Those variants are clinically important, and the model has never seen them. Supporting them is an architecture change, not a bug fix.
One genuine bug, found by writing this report54 rows record the target as "Wild Type" rather than "WT". The check matched only the literal "WT", so these went down the mutation path, failed to parse, and were discarded under a reason that was false. They are wild type. Fixed; the count is small but the drop reason was actively misleading.

Single-shot inhibition data, traced to the optimiser

Percent-inhibition measurements at a stated concentration are not potencies, and an earlier version of this project ignored them. They are converted to a potency bound and enter training as censored records. The reason to trace them rather than assert them is that this project once lost 35,495 rows between the training table and the optimiser, 98.7 percent of the greater-than records, and it went unnoticed for two cycles.

Conversion ruleIC50 = C(100-I)/I with margins at 20 and 80 percent inhibition. Below 20 percent gives an upper bound on potency (class Low); above 80 percent a lower bound (class High); between them the record is uninformative and is dropped.
stagerows
KKB percent-inhibition records (all targets)449,460
in the training table after conversion107,787
surviving embedding coverage, i.e. reaching the optimiser107,784
lost to embedding coverage3

100.0 percent of the converted single-shot rows reach the optimiser. By class: {'Low': 106685, 'High': 1099}.

Generated by scripts/46_provenance_report.py from data/provenance_audit.json. Sequence identity is shown as an md5 of the full amino-acid sequence, so two rows carry the same protein if and only if their md5 matches.