Announcing a multi-round prediction league to forecast forthcoming RCT estimates
Decades of social science RCTs have generated a wealth of rigorous impact estimates. By carefully isolating cause and effect, they offer insight into what policy alternatives achieve – and, more generally, into the forces driving society and the economy.
But this body of evidence is complex and hard to use. Individual studies can be noisy and difficult to generalize, while traditional methods for synthesizing them into balanced policy advice are slow, expensive, quickly outdated, and lose context.
A better science is one that makes better predictions: the more accurately we can forecast the results of trials that haven't yet resolved, the more we can claim to know. Advances in AI point in exactly this direction, demonstrating dramatic capabilities to synthesize vast amounts of knowledge and generate nuanced predictions. Indeed, recent work suggests that large language models can forecast the effects of survey experiments no worse than human forecasters, and that purpose-built, fine-tuned methods forecast causal effect sizes more accurately than off-the-shelf models – all indications that the approach has real headroom. But many challenges remain: for example, models trained on the published literature inherit publication bias, and machine prediction introduces a trade-off between accuracy and interpretability.
We aim to accelerate research toward a future where policy choices are guided by the best available evidence, and where new experiments are designed to maximize collective learning. To this end, we are announcing a prediction league: an open contest in which forecasting methods are scored against experimental results they have never seen.
As a first step, we are launching a pilot prize in which contestants train on a curated set of open access literature.
$1M+ pool committed across multiple rounds
There are two divisions teams can participate in.
No restrictions on model type. Any approach is permitted.
The purpose of this division is to establish a ceiling on performance of unrestricted models.
Must provide a self-contained package that regenerates the end-to-end pipeline, including feature construction from training data, model training, and generating forecasts on test data.
Reproducible submissions compete in both divisions.
Each division will award multiple prizes; prize recipients are selected based on objective, nondiscriminatory criteria related to scientific achievement in responsible AI development.
We're aiming to launch in fall 2026; the timeline below outlines our current best estimates.
The tournament is open to anyone who wants to forecast experimental results, whether entering alone or as a team of any size.
Entrants may come from any background, including but not limited to academic researchers, industry ML teams, independent forecasters, and evidence-synthesis teams. No specific disciplinary background is required. Entrants must certify that their models or research outputs submitted will be used for noncommercial purposes only. Entry is open globally, except where participation or receipt of prize funds is restricted by applicable law.
If a participant is an author of a forthcoming trial, or otherwise in a position to know results before they are reported, we ask that you disclose this during registration. Such conflicts will be assessed case by case.
The prediction set is a collection of ongoing, unresolved social science RCTs. For each RCT, participants receive its pre-analysis plan alongside a set of forecasting questions, submitting a predicted mean and variance for each study outcome.
Any study that resolves or shares results before the submission deadline will be removed from the scoring set.
For each trial, forecasts are scored on both the accuracy of the point estimate and the calibration of the uncertainty bounds, against the real effect once the trial resolves. Both divisions are scored by the same rule – the reproducible division adds the requirement that a winning submission include a fully reproducible package. A submission that meets the reproducible standard competes for both prizes. The full scoring and reproducibility criteria will be set out in the participant terms.
Participants retain ownership of their methods and code. The tournament makes no claim to either.
As a condition of claiming a prize, a winning submission provides us with a license to use and share the submission, alongside a detailed methodology note, so that others can learn from and build on it.
Yes. The tournament is a research contribution as well as a competition, and the pilot's findings will be published.
Winning methods are also published. As a condition of claiming a prize, winners will submit documentation of their approach and provide a license to use and share a copy of their model. The exact timeline and format are set out in the participant guidelines.
Registration will open in Fall 2026; in the meantime, register your interest through the Register button below to receive updates.
Yes. This is the first round, and further rounds are intended to follow, building on what we learn from the pilot. We are currently working to unlock a larger body of scientific literature for future rounds.
We'll be in touch as the tournament firms up. Thanks for your interest.
Something went wrong. Please try again, or email predictionleague@agency.fund.