Upscale Logo
Financial MarketsOctober 5

Can AI Pass a Prop Firm Challenge? Alpha Arena Results

Stanislav
StanislavTrading Research Lead
Can AI Pass a Prop Firm Challenge? Alpha Arena Results

Can an AI trade for you? The clearest real-world test took place in fall 2025, in the Alpha Arena experiment. Six of the largest language models — GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, DeepSeek V3.1 and Qwen3 Max — each received $10,000 in real money and traded crypto for 17 days with no human involvement. Four models finished with losses of 42% to 59%. The two profitable ones, Qwen3 Max (+22.3%) and DeepSeek V3.1 (+4.9%), doubled their accounts in the first ten days, then gave back more than half of their gains within days. By our estimate, on a prop challenge those days would have breached the standard daily drawdown limit, and the account would have been closed along with all of its profit. Below, we break down what happened, what the research says and what traders should take away from it.


This is a fact-based breakdown: the experiment's numbers, academic research and our own calculation under challenge rules. For how to use AI in your own trading, see our guide "How to Trade with AI".

What is Alpha Arena

1.png

Alpha Arena is an experiment run by the US research company Nof1. It tested a simple question: can a general-purpose language model — the same kind you chat with — trade a real market on its own? The models didn't trade on historical data or in a simulator. They traded live money, and every mistake cost real dollars.

Experiment setup

ParameterSetup
ParticipantsGPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, DeepSeek V3.1, Qwen3 Max
Capital$10,000 in real money per model
MarketPerpetual futures (perps) on the decentralized exchange Hyperliquid
AssetsBTC, ETH, SOL, XRP, DOGE, BNB
PeriodOctober 18 — November 3, 2025
Human roleNone: the models decided what to buy, which leverage to use and when to exit

All models received the same instructions and the same data: prices, indicators and the state of their own account. They couldn't go online or read the news. The only difference was the decisions each model made.

How this differs from a regular test

Most tests of AI in trading are backtests: the model "trades" on historical data. The authors of the LiveTradeBench study point to the main problem with this approach: information leakage. The model may have seen that data during training, in which case it isn't forecasting — it's remembering. Alpha Arena had no such shortcut: the models traded in real time, in a market whose future nobody knew, and every trade was publicly visible.

Season 1 results

2.png

The season ended on November 3, 2025. The final results were published by Nof1's founder and reported by ForkLog:

PlaceModelFinal balanceResult
1Qwen3 Max (Alibaba)$12,231+22.3%
2DeepSeek V3.1$10,489+4.9%
3Claude Sonnet 4.5 (Anthropic)$5,799−42.0%
4Gemini 2.5 Pro (Google)$5,445−45.6%
5Grok 4 (xAI)$4,208−57.9%
6GPT-5 (OpenAI)$4,126−58.7%

The official Season 1 pages on Nof1's website are no longer available, so we rely on the figures published by the project's founder. Media reports quote slightly different values, for GPT-5 for example, but that doesn't change the ranking or the overall conclusion.

Two of the six models made money. The other four lost nearly half of their account or more. But the table only shows the finish line — the real story is the path to it.

The winners: to the top and back

By October 27, the tenth day of the season, both future winners had doubled their accounts. According to the South China Morning Post, DeepSeek grew to $22,900 (+126%) and Qwen to $20,850 (+108%).

Then the market took almost all of it back. Over the following week, DeepSeek lost more than half of its balance, and Qwen fell 41% from its peak.

ModelBalance on October 27Finish on November 3Drop from peak
DeepSeek V3.1$22,900$10,489−54%
Qwen3 Max$20,850$12,231−41%

The winners had very different styles. PANews analyzed the models' trades over the first nine days. DeepSeek traded like a patient trend follower: only 17 trades, an average of 49 hours in a position, small stops and large profit targets. Qwen acted like a gambler: it usually held a single position, bet big on it and often used 25× leverage — the maximum allowed in the competition.

Simple math shows how dangerous that leverage is. At 25×, a 4% price move against the position wipes out 100% of the margin in it. In crypto, a 4% move is a perfectly ordinary day. While the market rose, leverage multiplied the profit. When the market turned, it multiplied the losses.

So the "winner" of Alpha Arena is the model that gave back less of its gains than the others. The final return hides what matters most: how deep the account sank along the way. For a prop trader, that is everything. We'll come back to it in the calculation below.

The losers: too many trades

The underperformers had a different problem, and a surprising one: they weren't taking too much risk. According to PANews, GPT-5 and Gemini used average leverage of less than 1× per position. What hurt them was frequency.

By October 27, Gemini had made 165 trades and spent more than $1,000 on fees alone. It exited at the slightest profit and the slightest loss, and each trade chipped away at the account. GPT-5 made 63 trades, and only about 20% of them were profitable. Both often entered the market later than the leaders, after the move had already happened.

The result: two opposite mistakes that every trader knows. The leaders were let down by excessive leverage, the underperformers by overtrading. That's why we keep telling traders that high leverage is mathematically incompatible with drawdown limits, and frequent trades without an edge only feed fees. Leverage of around 5× is a deliberate choice: at that level, one bad trade can't wipe out the account.

The "character" of AI models

Summing up the season, Nof1 founder Jay Azhang noted that the models showed consistent tendencies that persisted throughout the season and across many changes to their instructions. In his words, it resembled an investing "personality."

3.png

In practice, each model repeated its style over and over, even when it stopped working. Qwen kept putting everything on a single highly leveraged position, while Gemini kept opening dozens of small trades. It's the same behavior as a trader who repeats the same mistake time after time.

Season 2: a different market, a different leader

In late 2025, Nof1 ran Season 1.5. Eight models traded US stocks this time, in four parallel competitions. According to the published results, first place went to a different model, not the Season 1 winner. That's one more argument: success in one short experiment doesn't automatically carry over to another market and another period.

Would AI models pass a prop challenge: our calculation

The final returns in Alpha Arena answer the question "who earned the most." A trader needs the answer to a different question: could the model trade on an account with risk rules? To find out, we overlaid the models' results on Upscale's challenge rules. These rules are typical of prop trading, so the conclusion holds beyond our platform.

The rules we check

FormatDaily limitOverall drawdown limit
Basic5% of the balance at the start of the day10% of the starting balance
Accelerated3% of the balance at the start of the day6% of the starting balance
Turbodoes not apply6% of the maximum balance

Two important details. The balance includes open positions: a loss that hasn't been closed yet still reduces the balance. And the trading day starts at 00:00 UTC, with the daily limit reset every day. Learn more in the challenge rules.

Results by model

ModelOverall limit (Basic, Accelerated)Daily limit (Basic, Accelerated)Limit from peak (Turbo)
Qwen3 Maxmost likely not breachedbreachedbreached: −41% from peak
DeepSeek V3.1most likely not breachedbreachedbreached: −54% from peak
Claude, Gemini, Grok, GPT-5breached: finished below $9,000no longer relevantbreached

The four losers are simple: an account that lost 42–59% would have been closed under any rule. The two profitable models are more interesting. They most likely wouldn't have breached the overall limit from the starting balance: they were in profit during the first days and finished above $10,000. What would have stopped them is the daily limit.

How we calculated it

4.png

A minute-by-minute balance history isn't publicly available, so we relied on published values: the peak on October 27 and the finish on November 3. That's enough to reach a conclusion. Let's use DeepSeek as an example.

Suppose that from October 28 to November 3, DeepSeek never lost more than 5% in a day. Then over those seven days its balance could have fallen to no less than 70% of its level at the start of October 28 (0.95 to the seventh power ≈ 0.70). The model finished with $10,489, which means that at the start of October 28 it would have had no more than $15,020.

But on October 27, DeepSeek's balance was $22,900. That means it would have had to lose about 34% in a single day, October 27. Either way, the daily limit is breached. The same calculation for Qwen gives a drop of at least 16% in one day.

Why a doubled account wouldn't have saved them

It seems logical: if a model doubled its account, it has a big cushion. But the daily limit is checked every day, not just at the start. Once the balance had grown to $22,900, 5% of it was about $1,145. The models were losing thousands of dollars a day.

When a limit is breached, the account is closed immediately and all accumulated profit is cancelled. The models never withdrew anything, so they would have lost all their earnings at once. If a trader on a funded account had requested a payout closer to the peak, the withdrawn part would have been kept. But an AI model doesn't think about that on its own: it's a human decision.

That's the point of the daily limit. It doesn't punish the trader — it stops them on the day when everything is going wrong, before a bad day turns into losing the whole account.

Limitations of the calculation

This is an estimate, not an exact simulation. We overlay the balance path on the risk limits, but we don't go through the challenge stages: in a real challenge, positions are closed when moving to the next stage, and the trading that followed would have played out differently. We also only have separate balance snapshots, not the full minute-by-minute history. The conclusion itself doesn't change: days with losses like these are incompatible with a 5% daily limit.

What the research says

Alpha Arena is a single 17-day experiment. Researchers have tested AI traders for longer periods and on other markets. Here are the key studies from the past year and their findings in plain language.

StudyWhat was testedWhat it found
StockBench (ICLR 2026)Trading Dow Jones stocks, March–July 2025Most AI agents failed to beat a simple buy-and-hold strategy. High scores on financial tests don't translate into profitable trading
FINSABER (KDD 2026)20 years of history, more than 100 stocksOver the long run, the edge of AI strategies disappears: they are too cautious in rising markets and too aggressive in falling ones
LiveTradeBench (University of Illinois)50 days of live trading, 21 modelsA model's place in general "intelligence" rankings has almost no connection to its trading results
Agent Market ArenaLive trading of crypto and stocksAI agents can trade profitably and beat buy-and-hold, but the result depends more on how the agent is built than on which model it uses
AlphaForgeBenchConsistency of decisionsThe same model under identical conditions produced noticeably different results from run to run

The overall picture: AI can trade profitably in certain periods, but a consistent edge over the market has not been proven yet. And where agents did perform well, what mattered most wasn't the model's "intelligence" but the rules it operated under.

Why a "smart" model isn't necessarily a good trader

A language model learned from text: it can explain leverage or a trend very well, but that doesn't mean it can manage risk. Nof1's founder noted that language models don't handle numerical time series well — and a price chart is exactly that kind of series.

There's a second problem too. A model answers the same question differently from one time to the next, and in trading that randomness turns into random trades. A human can develop discipline; a model follows exactly the rules it has been given — and nothing beyond them.

What this means for traders

5.png

Can AI trade for you

Technically, yes: an AI model can already open and close trades without a human. But so far there's no evidence that it can do so consistently profitably while staying within risk limits. The experiments show the opposite: models make money during a lucky stretch and give it back when the market turns.

That doesn't mean AI is useless in trading. It works well as an assistant: it breaks down news, helps write and check code and analyzes your trades. We covered five levels of using AI in trading — from analyst to autonomous agent — in our guide "How to Trade with AI".

Risk limits belong in the code, not in the model's instructions

If you automate your trading, the main lesson of Alpha Arena is simple: you can't count on the model to stop by itself. Maximum leverage, position size and a daily stop must be hard-coded as rules in a program the AI can't bypass. We explained how to connect a bot to a challenge and which rules apply to automated trading in our article "API Trading".

Humans make the same mistakes

Excessive leverage, overtrading, trying to win back a loss — these mistakes aren't unique to AI. If your challenge has closed, the AI Challenge Report in your dashboard will show statistics across all your trades, your three main mistakes and an improvement plan for each of them. It's a good example of how AI is useful in trading: not to predict the market, but to help a person see their own mistakes.

Where to start if you want to trade with AI

Start with tasks where the AI's mistakes cost nothing: ask it to explain an unfamiliar term, break down a news story or analyze an export of your past trades. The next step is code. AI can help you write an indicator or a simple bot, but run the first versions only on a demo account, and never paste your API key into the chat. Move to a real challenge only once the strategy has proven itself on demo and the limits on leverage, position size and daily loss are hard-coded. Each of these steps is covered in detail in our guide "How to Trade with AI".

A short experiment doesn't prove that an AI model can trade. Over 17 days of Alpha Arena, two of the six models made money, but both gave back more than half of their profit along the way. In the second season, on a different market, the leader changed. Academic research confirms the same picture: AI has no consistent edge over the market yet.

What matters isn't the model's "intelligence" but the rules it trades by. The models lost money not because they lacked knowledge, but because of risk management mistakes: the leaders were let down by high leverage, the underperformers by overtrading and fees. Where AI agents did perform well, what mattered most was how the agent was built and the risk limits built into it.

For a trader on a challenge, the takeaway is practical. Use AI as an assistant — for analysis, code and reviewing your trades — not as an autopilot. If you automate your trading, limits on leverage, position size and daily loss must be hard-coded. The daily drawdown limit works the same way: it stops trading on a bad day, before a lost day turns into a lost account.


Start now: 👉 Upscale.trade | Telegram bot

Follow us: 📺 YouTube | 𝕏 Twitter

Community: 💬 Telegram chat | 🎮 Discord

What is Alpha Arena?

Alpha Arena is an experiment by the research company Nof1 in which language models traded real money with no human involvement. In Season 1, from October 18 to November 3, 2025, six models — GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, DeepSeek V3.1 and Qwen3 Max — each received $10,000 and traded crypto perpetual futures on the Hyperliquid exchange. All models received the same instructions and data.

Which AI won Alpha Arena?

Season 1 was won by Alibaba's Qwen3 Max, which finished with $12,231, or +22.3%. DeepSeek V3.1 came second with +4.9%. The other four models ended in the red: Claude Sonnet 4.5 (−42.0%), Gemini 2.5 Pro (−45.6%), Grok 4 (−57.9%) and GPT-5 (−58.7%). In the second season, on US stocks, a different model won.

Why did most AI models lose money?

For the underperformers, the main problem was overtrading. According to PANews, Gemini made 165 trades in nine days and spent more than $1,000 on fees, while only about 20% of GPT-5's trades were profitable. Their leverage was low. The leaders, on the other hand, were let down by high leverage: Qwen often traded at 25×, and after the market turned, both profitable models gave back more than half of their gains.

Can an AI pass a prop challenge?

There's no evidence of that yet. By our estimate, even the two profitable Alpha Arena models would have breached the standard 5% daily drawdown limit: on certain days they lost 15% of their balance or more. When a limit is breached, the account is closed and all accumulated profit is cancelled. The four losing models would also have breached the overall drawdown limit.

Can you copy AI models' trades?

It's a risky idea. The result of a short experiment doesn't guarantee future profit: in the second season of Alpha Arena the leader changed, and academic research finds no consistent edge for AI traders over the market. PANews analysts also noted that luck played a part in the leaders' success: they happened to align with the overall market rise. On top of that, trades with leverage of up to 25× are incompatible with the drawdown limits of most prop challenges.

Which AI is best at crypto trading?

There's no definitive answer. Qwen3 Max performed best in Season 1 of Alpha Arena, but 17 days is too short to draw conclusions. The LiveTradeBench study showed that a model's place in general rankings has almost no connection to its trading results, and Agent Market Arena found that how the agent is built affects the result more than the choice of model.

Can you trade with AI on Upscale?

Yes. On Upscale you can connect your own bot via API, including one written with the help of AI — see the details in our article "API Trading". Automated trading is subject to the same drawdown limits as manual trading. And after a challenge closes, you can request an AI Challenge Report: trade statistics, your three main mistakes and an improvement plan.

Ready for funded capital?

Pass the challenge, get up to $200,000

Suggested posts