Trusting stored labels produces a comparatively easy combined condition target. The score cannot be read as prospective laboratory performance.
Computational chemistry · transparent reproduction
When cleaning the data changes the answer.
Reaction prediction can look more accurate when the dataset quietly makes the task easier. This lab separates role assignment, rare-component policy, duplication, and split overlap—then shows exactly which conclusion each decision supports.
- Source reactions
- 1,771,032
- Published score range
- 35–68%
- Current evidence
- Stable v1 audit
Same model family. Different definition of the task.
Evidence status · research product v1.0.0
Baselines reproduce. Similarity limits are measured.
The checksum-pinned four-variant archive reproduces every peer-reviewed baseline within 0.46 percentage points. Exact reaction keys do not cross the split, but 5.78% of test products and 80.84% of nonempty test scaffolds occur in training. Model scores remain published references because exact checkpoint and prediction bundles are not archived.
Interactive 2 × 2 audit
The accuracy inflation microscope
Change one decision at a time. The dataset—not the neural architecture—changes.
Prespecified v1 secondary audit
Exact separation is not chemical novelty.
The full split and a frozen 1,000-row similarity sample expose what the identity check cannot see.
3,784 of 65,444 valid test products also occur canonically in training.
51,617 of 63,852 nonempty test scaffolds occur in training.
Every test row's source-file category is represented in training; this is not a patent-family identifier.
Select a frozen Morgan/Tanimoto threshold
A majority of the prespecified sample has a moderately similar training product. This describes product representation overlap, not reaction equivalence.
Official cleaning log · reaction-string / delete-rare variant
Cleaning is a sequence of estimand changes.
Select a stage to inspect what was removed and what that operation does not prove.
Stage 01
Source extraction
1,771,032 reactions enterThe official USPTO-derived ORD extraction before the condition-benchmark filters.
Boundary: A large source count says nothing about independence, coverage, or label quality.
Read the metric correctly
Three numbers. Three different questions.
Exact match
Did the predicted set exactly match the recorded solvents and agents? Partial chemical usefulness is not measured.
Top-3
Was the exact recorded combination among the model's three highest-ranked combinations? It is not top-1 accuracy.
Baseline
How well does a frequency-informed guess perform? A high model score matters less if the task is already predictable from prevalence.
Stop at the evidence boundary
What this study can—and cannot—say.
Supported now
- The complete four-variant archive matches its official checksum.
- All four frequency baselines reproduce within 0.46 percentage points.
- No exact declared reaction input key crosses from training into test.
- Product identity, scaffold, source-file, date, and sampled similarity overlap are quantified.
Not supported
- That the four neural-model cells reproduce locally.
- That source-file overlap is patent-family leakage.
- That any score represents prospective wet-lab success.
- That observed benchmark ease implies research misconduct.
Primary evidence