Reaction Integrity Lab

Computational chemistry · transparent reproduction

When cleaning the data changes the answer.

Reaction prediction can look more accurate when the dataset quietly makes the task easier. This lab separates role assignment, rare-component policy, duplication, and split overlap—then shows exactly which conclusion each decision supports.

Source reactions
1,771,032
Published score range
35–68%
Current evidence
Stable v1 audit
INPUT REPRESENTATIONOUTPUT LABELS

Same model family. Different definition of the task.

Evidence status · research product v1.0.0

Baselines reproduce. Similarity limits are measured.

The checksum-pinned four-variant archive reproduces every peer-reviewed baseline within 0.46 percentage points. Exact reaction keys do not cross the split, but 5.78% of test products and 80.84% of nonempty test scaffolds occur in training. Model scores remain published references because exact checkpoint and prediction bundles are not archived.

Interactive 2 × 2 audit

The accuracy inflation microscope

Change one decision at a time. The dataset—not the neural architecture—changes.

How are reaction roles assigned?
What happens to rare components?
Reproduced baseline · published modelA
Frequency baseline 52%
Model top-3 exact match 67%
Gain over baseline
+15 pp
Normalized improvement
32%
Evidence owner
Wigh et al. (2024)

Trusting stored labels produces a comparatively easy combined condition target. The score cannot be read as prospective laboratory performance.

Prespecified v1 secondary audit

Exact separation is not chemical novelty.

The full split and a frozen 1,000-row similarity sample expose what the identity check cannot see.

Product identity5.78%

3,784 of 65,444 valid test products also occur canonically in training.

Product scaffold80.84%

51,617 of 63,852 nonempty test scaffolds occur in training.

Source-file category100%

Every test row's source-file category is represented in training; this is not a patent-family identifier.

Select a frozen Morgan/Tanimoto threshold

Maximum product similarity≥ 0.70
60.5%
Sample count
605 / 1,000
Wilson 95% interval
57.44–63.48%
Fingerprint
Morgan r=2 · 2,048 bit

A majority of the prespecified sample has a moderately similar training product. This describes product representation overlap, not reaction equivalence.

Official cleaning log · reaction-string / delete-rare variant

Cleaning is a sequence of estimand changes.

Select a stage to inspect what was removed and what that operation does not prove.

Stage 01

Source extraction

1,771,032 reactions enter

The official USPTO-derived ORD extraction before the condition-benchmark filters.

Boundary: A large source count says nothing about independence, coverage, or label quality.

Read the metric correctly

Three numbers. Three different questions.

01

Exact match

Did the predicted set exactly match the recorded solvents and agents? Partial chemical usefulness is not measured.

02

Top-3

Was the exact recorded combination among the model's three highest-ranked combinations? It is not top-1 accuracy.

03

Baseline

How well does a frequency-informed guess perform? A high model score matters less if the task is already predictable from prevalence.

Stop at the evidence boundary

What this study can—and cannot—say.

Supported now

  • The complete four-variant archive matches its official checksum.
  • All four frequency baselines reproduce within 0.46 percentage points.
  • No exact declared reaction input key crosses from training into test.
  • Product identity, scaffold, source-file, date, and sampled similarity overlap are quantified.

Not supported

  • That the four neural-model cells reproduce locally.
  • That source-file overlap is patent-family leakage.
  • That any score represents prospective wet-lab success.
  • That observed benchmark ease implies research misconduct.

Primary evidence

Follow every number to its source.