Popper · falsifiable hypothesis compiler

Negative controls

The most important part of the experiment. An LLM is excellent at producing plausible narratives, so the compiler could "succeed" for the wrong reason. We run three controls on the same corpus so the only variable is the reasoning:

methodCompositeNoveltyGroundingTestabilityfalsifiable rate
compiler0.940.900.931.001.00
llm-only0.000.950.000.860.00
keyword0.460.530.780.290.00
random0.000.381.000.000.00

Read this as the falsification test of the whole project. If the compiler row is not clearly higher than the control rows on composite and grounding — and if the controls' falsifiable-rate is not far lower — then structured compilation is not adding what we claim, and we report that honestly.