Make money doing the work you believe in
You may have seen Warren’s reply on X, but here it is: Thanks for this, and fair enough. FutureEval deserves credit where credit is due. Running frontier bots against the community and your Pros on the same questions over time is exactly the kind of disciplined, public work this field needs.
Let me sharpen what I am actually after. The matchup I want to see is Good Judgment Superforecaster teams against the top models in an ongoing competition under the same conditions, same questions, same clock, same rules. That specific matchup is the one no one has done yet despite what headlines (“superforecasters vs AI”) claim. And we would include a mix of binomials, ordinals, and multinomials across a variety of timelines using real-world questions proposed by public and private decision makers, all of which would be displayed on a public dashboard.
And a quick correction on my end. The “whoever ran hottest last quarter” line was aimed at the self-proclaimed “superforecasters” who get anointed in the press after one hot streak, not at your Pros. But placed where it was, it reads as a swipe at how you select your Pros, and I can see why it landed that way. That’s on me, and I apologize for it.
Part of the confusion is that “Superforecaster” is sometimes used as a catch-all for anyone who claims to be forecasting well. It’s our trademark, and we need to be vigilant about its unauthorized use, which can be tougher to do with media but elicits a note from our counsel when used commercially.
