A multi-round prediction league aimed at spurring the development of AI-augmented models for forecasting the impacts of forthcoming social science RCTs
Decades of social science RCTs have built a rich evidence base, though it is unclear how we can turn that accumulated knowledge into predictions about new interventions.
This initiative aims to answer a key question: to what degree can AI-augmented systems forecast the direction and magnitude of treatment effects before a trial completes?
Recent work suggests large language models can rank the effects of social science survey experiments comparable to pooled human forecasters, and that purpose-built, fine-tuned methods forecast causal effect sizes more accurately than off-the-shelf models – evidence that the approach has real headroom.
However, such approaches tend to overestimate effect sizes, and variance increases with unfamiliar designs, contexts, and outcomes. Forecasting a trial's result from its design, context, and measurement strategy remains difficult due to multiple sources of uncertainty, and a persistent tradeoff between accuracy and interpretability, with limited application so far to field experiments.
By rewarding specialised methods under a proper scoring rule and setup, the prediction league aims to establish whether such performance holds prospectively. We are launching this as a pilot progress prize and expect it to serve as a proof of concept for continued expansion.
$500k+ in prizes for the pilot.
There are two divisions teams can participate in.
No restrictions on model type. Any approach is permitted.
The purpose of this division is to establish a ceiling on performance of unrestricted models.
Must provide a self-contained package that regenerates the end-to-end pipeline using only competition provided or publicly available materials; use of pre-computed features of unknown origin disqualifies.
Reproducible submissions compete in both divisions.
Each division will award multiple prizes; prize recipients are selected based on objective, nondiscriminatory criteria related to scientific achievement in responsible AI development.
We're aiming to launch in fall 2026; the timeline below outlines our current best estimates.
The tournament is open to anyone who wants to forecast experimental results, whether entering alone or as a team of any size.
Entrants may come from any background, including but not limited to academic researchers, industry ML teams, independent forecasters, and evidence-synthesis teams. No specific disciplinary background is required. Entrants must certify that their models or research outputs submitted will be used for noncommercial purposes only. Entry is open globally, except where participation or receipt of prize funds is restricted by applicable law.
If a participant is an author of a forthcoming trial, or otherwise in a position to know results before they are reported, we ask that you disclose this during registration. Such conflicts will be assessed case by case.
The prediction set is a collection of ongoing, unresolved social science RCTs. For each RCT, participants receive its pre-analysis plan alongside a set of forecasting questions, submitting a predicted mean and variance for each study outcome.
Any study that resolves or shares results before the submission deadline will be removed from the scoring set.
For each trial, forecasts are scored on both the accuracy of the point estimate and the calibration of the uncertainty bounds, against the real effect once the trial resolves. Both divisions are scored by the same rule – the reproducible division adds the requirement that a winning submission include a fully reproducible package. A submission that meets the reproducible standard competes for both prizes. The full scoring and reproducibility criteria will be set out in the participant terms.
Participants retain ownership of their methods and code. The tournament makes no claim to either.
As a condition of claiming a prize, a winning submission provides us with a license to use and share the submission, alongside a detailed methodology note, so that others can learn from and build on it.
Yes. The tournament is a research contribution as well as a competition, and the pilot's findings will be published.
Winning methods are also published. As a condition of claiming a prize, winners will submit documentation of their approach shortly after results are announced and provide a license to use and share a copy of their model. The exact timeline and format are set out in the participant guidelines.
Registration will open in fall 2026; in the meantime, register your interest through the Register button below to receive updates.
Yes. This is the first round, and further rounds are intended to follow, building on what we learn from the pilot.
We'll be in touch as the tournament firms up. Thanks for your interest.