Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets

We introduce Prediction Arena, a benchmark for evaluating AI models'predictive accuracy and decision-making by enabling them to trade autonomously on live prediction markets with real capital. Unlike synthetic benchmarks, Prediction Arena tests models in environments where trades execute on actual exchanges (Kalshi and Polymarket), providing objective ground truth that cannot be gamed or overfitted. Each model operates as an independent agent starting with $10,000, making autonomous decisions every 15-45 minutes. Over a 57-day longitudinal evaluation (January 12 to March 9, 2026), we track two cohorts: six frontier models in live trading (Cohort 1, full period) and four next-generation models in paper trading (Cohort 2, 3-day preliminary). For Cohort 1, final Kalshi returns range from -16.0% to -30.8%. Our analysis identifies a clear performance hierarchy: initial prediction accuracy and the ability to capitalize on correct predictions are the main drivers, while research volume shows no correlation with outcomes. A striking cross-platform contrast emerges from parallel Polymarket live trading: Cohort 1 models averaged only -1.1% on Polymarket vs. -22.6% on Kalshi, with grok-4-20-checkpoint achieving a 71.4% settlement win rate - the highest across any platform or cohort. gemini-3.1-pro-preview (Cohort 2), which executed zero trades on Kalshi, achieved +6.02% on Polymarket in 3 days - the best return of any model across either cohort - demonstrating that platform design has a profound effect on which models succeed. Beyond performance, we analyze computational efficiency (token usage, cycle time), settlement accuracy, exit patterns, and market preferences, providing a comprehensive view of how frontier models behave under real financial pressure.

Paper

References (17)

07in Financial Machine2018 · Advances
09Exit Selection Quality: Holding winners and cutting losers appropriately. Models that exit positions at the right time achieve better risk-adjusted returns
10Agent Decision-Making: Receive market context and make trading decisions through autonomous reasoning, using research tools and knowledge management capabilities
11though the underlying logic is identical to nettingRisk Management and Position Constraints
12Position Sizing When Uncertain: Appropriate position sizing limits losses when wrong. Models with lower max drawdowns manage risk more effectively

Scroll for more · 5 remaining

Similar papers

© 2026 NYSGPT2525 LLC