Verifying the Verifiers: Towards Autonomous Policy Evaluation

Olaf Willner and David Yanagizawa-Drott · Working paper, July 2026

Abstract

AI is now capable of producing policy evaluation papers, end to end, that look like real papers. But are they to be trusted? If they can be, the technology would open the door to evidence generation at unprecedented scale: cheap, rapid, and global. Checking thousands of papers by hand is infeasible, so verification itself must be automated — which raises the question of whether the verifiers can be trusted. Here, we study the reliability of automated verifiers powered by LLMs. We first develop a versioned taxonomy, the CRED codebook (Classification of Research Errors and Defects), followed by a benchmark and scoring method that quantify a verifier's ability to detect errors under that taxonomy. This lets us compare verifiers and trace the frontier between cost and detection recall over time. We find remarkable recent progress. Eighteen months ago, with the first ‘reasoning’ models, detection recall was around 30% — more than two of every three errors missed. By February 2026, the best models missed one in six. In July 2026, the first model reached 99% recall (Codex 5.6 Sol), suggesting a new era for verifiers. Costs have collapsed alongside: the cheapest price of 50% recall fell roughly 90 times over the past year, and the cheapest price of 20% recall fell more than 400 times in eight months. Ensembles built largely from models with open weights prove highly cost effective. Reliability, however, is not solved: recall degrades when a paper contains multiple errors, and confidence scores remain too poorly calibrated to decide when human review is unnecessary. We conclude by discussing what is verifiable in principle and what is not. Together, our results suggest that automated verification of policy evaluation research may soon be feasible at scale, reliably and cheaply.

Going to outer space with new space: The rise and consequences of evolving public-private partnerships

Avishai Melamed, Adi Rao, Olaf de Rohan Willner and Sarah Kreps · Space Policy, vol. 68 (2024)

Abstract

What explains the commercialization of key government space projects through the incorporation of New Space? The newer generation of private companies have seen a significant increase in government contracting as they become instrumental for national security missions and high-profile civil projects. The turn to New Space companies, particularly those entrepreneurially-driven and privately funded, deviates from the governments' historical reliance on more traditional private partners. We argue that the turn towards New Space was neither inevitable nor monocausal, but rather the product of the confluence of the upstart sector's cost-efficient service offerings and rising public profile, which coincided with a period of renewed international competition. New Space firms indeed distinguished themselves by offering affordable products, an innovative production process, and a unique brand of prestigious reputation otherwise unavailable at national programs and older aerospace companies. However, these services were only deemed necessary for integration with the public sector because of the heightened importance of security and national status amidst a perceived return to great power competition. A new generation of public-private partnerships offered the only strategy for spacefaring states to attain and maintain a competitive position in an environment where non-state actors can match government accomplishments and capabilities. However, the utility of current integrative policy threatens to globalize not only the strengths but also the weaknesses of New Space, undercutting the very goals their adoption was meant to achieve.