Method
Popper treats a hypothesis as a compiled artifact. A traditional compiler turns source code into an executable that must pass tests. Popper turns literature into a hypothesis that must pass validation.
The falsifiable hypothesis schema
A hypothesis is admitted only if it is structurally falsifiable. Missing an independent variable, a falsification condition, a quantitative prediction, or a source id is not a stylistic lapse — it is a compile error. Required fields:
- Independent / dependent variable — what you change, what you measure
- Population / system and measurement
- Expected outcome and an explicit falsification condition
- Mechanism — the causal chain linking IV to DV
- Quantitative prediction — direction + numeric magnitude window + unit + confidence
- Evidence — every claim tagged with a real
source_id
The four scores
- Grounding — fraction of evidence with a verifiable source id, discounted by an LLM audit that checks the source actually supports the claim (catches articulate autocomplete that cites a real id for an unsupported claim).
- Testability — objective structural falsifiability; how many required fields are present and non-trivial. Computed without the model.
- Novelty — distance from existing corpus claims, capped low if an LLM check finds the connection is already directly stated. 1.0 = no known direct connection.
- Expert agreement (rediscovery only) — semantic match to the removed ground-truth discovery.
Composite is the geometric mean of the applicable scores. A zero in any dimension tanks the composite — by design. A beautifully written but ungrounded hypothesis must not score well.
The benchmark
Rediscovery (Group A): real historical discoveries (CRISPR-Cas9, nucleoside-modified mRNA, CTLA-4 checkpoint blockade). We remove the discovery paper, hand the compiler only the prior literature, and ask whether it reconstructs the leap.
Attractive nonsense (Group B): historical dead ends (cold fusion, phlogiston) whose narratives fit the prior evidence. A trustworthy system must not rank these highly.
Controls: random traversal (noise floor), keyword co-occurrence (statistics baseline), and llm-only (no corpus). The compiler must beat all three on the same input, or the central claim fails.
Local-first
Every generation, audit, and score is produced by a 35B model running on the author's own hardware via an OpenAI-compatible endpoint — no cloud APIs, no keys. The whole point is a scientific instrument you can own, inspect, and rerun.