Negative controls
The most important part of the experiment. An LLM is excellent at producing plausible narratives, so the compiler could "succeed" for the wrong reason. We run three controls on the same corpus so the only variable is the reasoning:
- random — connect two random claims. Does the compiler beat noise?
- keyword — most frequent co-occurring concepts, templated. Does semantic reasoning beat co-occurrence statistics?
- llm-only — same model, same ask, but no corpus. Does the knowledge graph actually help?
- compiler — the real method: grounded, structured compilation.
| method | Composite | Novelty | Grounding | Testability | falsifiable rate |
|---|---|---|---|---|---|
| compiler | 0.94 | 0.90 | 0.93 | 1.00 | 1.00 |
| llm-only | 0.00 | 0.95 | 0.00 | 0.86 | 0.00 |
| keyword | 0.46 | 0.53 | 0.78 | 0.29 | 0.00 |
| random | 0.00 | 0.38 | 1.00 | 0.00 | 0.00 |
Read this as the falsification test of the whole project. If the compiler row is not clearly higher than the control rows on composite and grounding — and if the controls' falsifiable-rate is not far lower — then structured compilation is not adding what we claim, and we report that honestly.