Is the compiler falsifiable?
Earlier iterations proved a knowledge-graph compiler can generate and scale. The remaining question is harder: are its hypotheses scientifically meaningful, or only linguistically convincing? Popper stops trying to make the compiler smarter and starts making it falsifiable — every hypothesis carries variables, a mechanism, a quantitative prediction, and source-level provenance, and is scored against negative controls and a historical rediscovery benchmark.
The claim under test is not "AI can generate hypotheses" (everyone knows that). It is: structured knowledge compilation produces higher-quality, falsifiable scientific hypotheses than unstructured generation. If the compiler bar is not clearly above the three control bars, the claim fails — and that is a real, publishable result either way.
Rediscovery benchmark
For each historical discovery we removed the discovery paper and gave the compiler only the prior literature. Could it reconstruct the leap?
Adversarial (attractive nonsense)
A discovery system must avoid attractive nonsense. These are historical dead ends whose narratives fit the prior evidence well. We want their reconstruction match to be low.
3 hypotheses per method per domain · audit on · compiled locally.