Popper · falsifiable hypothesis compiler

Is the compiler falsifiable?

Earlier iterations proved a knowledge-graph compiler can generate and scale. The remaining question is harder: are its hypotheses scientifically meaningful, or only linguistically convincing? Popper stops trying to make the compiler smarter and starts making it falsifiable — every hypothesis carries variables, a mechanism, a quantitative prediction, and source-level provenance, and is scored against negative controls and a historical rediscovery benchmark.

compiler
0.94
mean composite
llm-only
0.00
mean composite
keyword
0.46
mean composite
random
0.00
mean composite

The claim under test is not "AI can generate hypotheses" (everyone knows that). It is: structured knowledge compilation produces higher-quality, falsifiable scientific hypotheses than unstructured generation. If the compiler bar is not clearly above the three control bars, the claim fails — and that is a real, publishable result either way.

Rediscovery benchmark

For each historical discovery we removed the discovery paper and gave the compiler only the prior literature. Could it reconstruct the leap?

CTLA-4 / PD-1 checkpoint blockade as cancer immunotherapy
held out: leach1996
1.00
best compiler rediscovery match
CRISPR-Cas9 as a programmable genome-editing tool
held out: jinek2012
0.00
best compiler rediscovery match
Nucleoside-modified mRNA as a safe, potent vaccine platform
held out: kariko2005
0.90
best compiler rediscovery match

Adversarial (attractive nonsense)

A discovery system must avoid attractive nonsense. These are historical dead ends whose narratives fit the prior evidence well. We want their reconstruction match to be low.

Room-temperature nuclear fusion in palladium electrodes (a known dead end)
0.50
compiler match (lower is better)
Phlogiston as the substance released during combustion (a known dead end)
0.00
compiler match (lower is better)

3 hypotheses per method per domain · audit on · compiled locally.