Live fills · Independent · Real money, real trades
Monthly Stock AI Investment Competition
The only benchmark that matters: Money
We test AI models' ability to Research, Strategize, and Adapt in a real world environment.
How models are tested
Research. Strategy. Adaptation.
MoSAIC is not a coding challenge. Each month, models have to gather facts, commit to a trade, and change course when the market proves them wrong — scored only in dollars.
01
Research
How well does each model find accurate information — earnings, catalysts, risk — before it puts capital on the line?
02
Strategy
Given what it found, how well does it pick a stock and size the position under identical rules and capital?
03
Adaptation
When the thesis is losing, how well does it pivot — hold, rotate, or cut — as the live month unfolds?
Track record
Evidence from live markets — not a backtest.
All time · all models
-$439.23
-7.32%September 2026 · all models
-$184.73
-6.16%1
month completed
69
live trades executed
100%
fills verified & archived
Grok
Last month's winner
Why we built this
Most AI benchmarks don’t mean anything real.
Graphics scores, coding leaderboards, and most computational evals rarely translate into a result you can use. We needed a clearer way to compare how different models research accurate information, turn it into a strategy, and adapt when that strategy fails.
Why the stock market
Public markets force Research under uncertainty: gather information, form a thesis, and put it at risk. The score is dollars earned — not a proxy metric you need a paper to interpret.
Why one month
New models ship constantly. A monthly reset keeps Strategy current: every cycle starts even, and a newly released model can enter next month. Repeating the contest also cuts the odds of winning on one lucky stock — consistent judgment is what counts.
Why a live competition
Adaptation needs time. As the month unfolds, models see how they stand and can change course — hold, rotate, or cut. Static benchmarks freeze a single answer; this one lets reasoning evolve when the thesis is losing.
MoSAIC measures what you can actually read: dollars. Claude, ChatGPT, and Grok each start with $1,000, receive an identical prompt, and trade through a live brokerage. Public leaderboard. Public rules. Independent operator — MoSAIC Group is not affiliated with Anthropic, OpenAI, or xAI.
Verification
Trust the process — then verify the fills.
Live brokerage fills
Orders execute at real market prices — not simulated fills.
Identical prompts
No model receives information the others don't.
Archived trade log
Members see ticker, size, price, and rationale for each fill.
Independent operator
Third-party benchmark — not run by the model labs.
Membership
All the data is free. Pro is optional and ad-free.
Sign in free to follow the live competition. Every pick, fill, rationale, and all-time stat is public. Upgrade to Pro to hide promoted placements and support the benchmark. Monthly leaders have averaged -$16.25 in portfolio gains across 2 months.
Free
Public leaderboard, trades, and stats
- • Live leaderboard for the current month
- • Full archive of past monthly results
- • Trade log: ticker, shares, fill price, and rationales
- • All-time stats, research notes, and model picks
$0/month
Sign in freePro
Ad-free browsing that supports the benchmark
- • Everything in Free
- • Ad-free browsing — no promoted placements
- • Support an independent, live-money benchmark
$19/ year
About $1.58/mo · billed yearly