Which company has the best AI model?
Liquidity-weighted aggregate sits at 31% across 8 Kalshi contracts.
Implied probability
Kalshi
31%
8 contracts
Polymarket
—
not bound
Cross-venue gap
—
single venue
24h move
+1pp
3h ago
24h volume
$20K
8 contracts
Closes
Dec 31, 2026
161 days
30-day trend
Bracket families
3 clusters across 8 contracts.
These contracts were grouped by title similarity. The headline aggregate combines all clusters; verify the cluster you actually need before quoting a number.
Heads-up — heterogeneous clusters
The top two clusters share only 30% of their title tokens — “Best AI on Dec 31, 2026” vs “Which AI company will have the best coding model on Dec 31, 2026”. The headline aggregate weights both, so the number on this page is meaningful only if the clusters resolve to the same question.
Cluster 1
Best AI on Dec 31, 2026
Cluster 2
Which AI company will have the best coding model on Dec 31, 2026
Which AI company will have the best coding model on Dec 31, 2026?: Anthropic
KXCODINGMODEL-26DEC-ANTH
Which AI company will have the best coding model on Dec 31, 2026?: OpenAI
KXCODINGMODEL-26DEC-OPEN
Which AI company will have the best coding model on Dec 31, 2026?: xAI
KXCODINGMODEL-26DEC-XAI
Cluster 3
Will Los Angeles D have the best record in Pro Baseball in the 2026 regular season
Analysis
This probability reflects market expectations about which AI company will have the best-performing model by December 31, 2026. The 30% price for ChatGPT contrasts with Claude's 62¢, suggesting traders believe Claude is the frontrunner based on recent benchmark performance and capability assessments. Market movements depend on concrete model releases, published benchmarks, and real-world capability demonstrations over the next five months. The resolution will likely depend on standardized evaluation metrics, head-to-head testing results, and consensus among AI researchers about which system achieves the strongest performance across reasoning, coding, and instruction-following tasks.
- ›Claude's market price of 62¢ versus ChatGPT's 14¢ reflects trader assessment of recent model versions and their performance on standard benchmarks through mid-2026
- ›New major model releases or significant capability improvements announced between July and December 2026 would directly move probabilities, as traders reassess relative performance
- ›The resolution criteria must specify how 'best' will be measured—whether by published benchmark scores, evaluator consensus, or specific capability demonstrations
- ›Trading volume varies substantially across contracts ($7,167 for ChatGPT versus $1,342 for Grok), suggesting liquidity disparities may amplify or dampen price movements
- ›Historical precedent shows AI capability rankings shift rapidly when new versions release, making the December deadline potentially volatile for repricing
What moved the line
- Jul 16Anthropic↓5pp59→54¢ · Kalshi
- Jul 19Anthropic↑5pp56→61¢ · Kalshi
- Jul 20Claude↑4pp62→66¢ · Kalshi
- Jul 21Claude↓4pp66→62¢ · Kalshi
- Jul 20Los Angeles D↑4pp64→68¢ · Kalshi
Recently closed in technology
- How many SpaceX Starship launches in 2026?last 24% · 1d
- Will OpenAI release GPT-5.6 before Jul 15, 2026last 97% · 15d
- Will SpaceX be added to the S&P 500 in Q2 2026last 4% · 22d
- Will SpaceX be assigned to Communication Services in the S&P-500last 95% · 34d
- Will OpenAI's valuation hit __ by December 31?: ↑$2.0Tlast 92% · 34d
These markets stopped trading. Last odds and any captured outcome are shown above — full settlement detail lives at the venue.
More like this
Adjacent prediction questions.
In technology
How we compute these odds
SimpleFunctions aggregates live prediction-market contracts from Kalshi and Polymarket. Each slug groups contracts that resolve on the same underlying event, identified by venue event_id.
For binary slugs, the headline probability is the liquidity-weighted mid-price across all bound contracts. For multi-outcome slugs (e.g. elections with 3+ candidates), the headline is the leader’s price; we never arithmetically average disjoint outcomes — that would produce a number with no real-world meaning.
Snapshots refresh every 5 minutes during market hours; daily aggregates are computed at 04:00 UTC. The 30-day sparkline is drawn from per-ticker daily means stored in market_indicator_daily; 24h delta and movement events are derived from the same source.
Last updated on this page: just now.