Claude vs ChatGPT vs Gemini for Trading: What a Real-Money Test Found

Last updated: October 5, 2026
By TradingBotsSimplified

No model can be called the best trader from the first real-money comparison of Claude, ChatGPT and Gemini at retail brokerages. In Condor Capital’s three-day test, Claude returned 2.9%, Gemini 0.9% and ChatGPT −0.7% on separate $500 accounts. But each model traded through a different broker, the sample lasted only three days, and the agents chose radically different activity levels.

The more durable finding is that broker tools and risk controls mattered at least as much as the model. ChatGPT completed 28 round trips, Claude 10 and Gemini 3. A useful AI trading agent should therefore be judged on instruction-following, permissions, execution quality, auditability and repeatability—not a tiny return leaderboard. That conclusion complements a separate 2026 live test of AI trading bots, where performance also deteriorated as testing moved closer to real execution.

What the real-money test actually measured

Condor Capital’s Digital Advice Report used funded taxable accounts and routed every order to a live market. It ran three separate experiments: a long-running 60/40 portfolio mandate, a 19-task capability matrix, and an open-ended trading challenge.

The headline returns came from the open-ended challenge. Each agent started with $500, could buy long stocks and ETFs, could not risk more than the account held, and had to return to 100% cash by the close. The table reports results for August 3–5, 2026. Claude traded at Robinhood, ChatGPT at Public and Gemini at Webull.

That design shows what one model-broker combination did under one prompt. It does not isolate model quality: changing the model also changed the brokerage, available tools, execution interface and check-in cadence. Condor also warns that the products evolve quickly and the agents are non-deterministic, so repeating the same instruction can produce a different path.

Claude led the three-day return—but this was not a model ranking

The results were real, but the experiment was too short and structurally confounded to establish a winner. Claude finished ahead, Gemini made a smaller gain and ChatGPT lost money. The spread from first to last was 3.6 percentage points.

Model and brokerEnding valueRealized P&LReturnRound tripsChosen cadence
Claude at Robinhood$514.34+$14.34+2.9%1010 checks per day
ChatGPT at Public$496.69−$3.31−0.7%28About every 5 minutes
Gemini at Webull$504.42+$4.42+0.9%35–8 checks per day

Source: Condor Capital’s report. The 3.6-point return range and cadence comparisons are TradingBotsSimplified calculations from the published figures.

All three agents had their best day on Tuesday and their worst day on Wednesday; Gemini stayed barely positive on Wednesday by trading very little. Condor reasonably interpreted that common pattern as evidence that the market tape influenced the result more than model identity. Three days also cannot establish a stable Sharpe ratio, drawdown profile or ability to generalize across market regimes.

Trade cadence explained more than the leaderboard

The prompt—“make as much money as you can”—specified an objective, not a strategy. Each agent filled in the missing rules differently. ChatGPT checked roughly every five minutes and completed 28 round trips, 2.8 times Claude’s activity and about 9.3 times Gemini’s. Claude used stops, scheduled fewer reviews and wrote lessons for its next session. Gemini traded only Nvidia, once per day, and did not set a stop.

That variation is the central engineering lesson. An LLM will supply its own interpretation when the specification omits the asset universe, signal, entry, exit, position sizing, order type, trading frequency or risk limit. More intelligence does not make those hidden choices valid.

A fair model comparison should freeze those decisions first. Use the same broker, tools, market data, clock, risk budget and deterministic order rules; repeat the trial because model output varies; and compare intended orders with actual fills. The familiar progression from backtesting to paper and live trading still applies when natural language replaces some of the code.

The broker mattered as much as the model

Condor’s 19-task matrix compared Claude and ChatGPT at Robinhood, Public and Webull. Each completed task earned two points, a workaround earned one and an unavailable task earned zero. Both models received the same score at each brokerage, which means the differences came from the broker interfaces rather than the model.

Broker interface in the testDerived scoreDerived completion rateMaterial gaps during the test
Robinhood31/3881.6%No native bracket order, crypto or short selling
Public33/3886.8%No watchlist tool, true GTC or native bracket order
Webull37/3897.4%Attached stop required a workaround

Method: TradingBotsSimplified summed Condor’s published task scores for either model at each broker and divided by the 38-point maximum. These are capability-completion rates, not safety or profitability ratings.

The snapshot can become outdated as brokers add tools. Verify the current interface before trading. Robinhood’s official Agentic Trading announcement describes a dedicated account and kill switch, while Public’s MCP disclosure says third-party agents can view data and place trades in the connected account. For current product-specific detail, see the separate guides to Robinhood’s AI trading agent and Webull Cloud MCP.

The most important failures were not losing trades

Condor documented Gemini inventing a whole-share limitation at Webull, then discovering a real $5 minimum fractional-order constraint only after an attempted trade. It also cited the old pattern-day-trader rule even though FINRA’s replacement rule had taken effect at the rule level.FINRA Regulatory Notice 26-10 confirms that the replacement rule became effective on June 4, 2026, but also allows member firms to phase in implementation through October 20, 2027. The old day-trade-count designation and $25,000 minimum may therefore still appear at a broker during its transition.

Those examples matter because an agent can be operationally fluent while factually stale. A wrong constraint may block a valid trade; a missing constraint may permit an invalid or dangerous one. Current broker rules, fees, tax treatment and regulations should therefore come from authoritative live sources, not the model’s memory.

Data and authority are separate risks. Robinhood states that data shared with a third-party provider leaves Robinhood’s security environment, and Public says it does not supervise or audit a connected third-party agent. Practical controls should include a ring-fenced account, the smallest required permissions, read-only or manual-approval mode during testing, symbol and notional limits, maximum daily loss, fill monitoring and an immediate disconnect path.

How to evaluate an AI trading agent for your own workflow

1. Define one narrow job. “Rebalance to fixed target weights” is testable; “make as much money as possible” lets the model invent a strategy.

2. Hold the environment constant. Compare models through the same broker, tools, data timestamps, account type and order permissions. Otherwise the broker becomes a hidden variable.

3. Repeat identical trials. Run the same frozen instruction across multiple dates and seeds. Record disagreement, refusals and interventions, not just returns.

4. Score operations before P&L. Measure instruction accuracy, order-preview accuracy, rejected orders, duplicate orders, slippage, fees, turnover, exposure, maximum drawdown and human intervention rate.

5. Use deterministic risk gates. Let the agent research or propose, but make code enforce the asset allowlist, position size, order type, price bounds, trading hours, daily loss and flattening rules.

6. Escalate capital slowly. Begin read-only, move to shadow orders, paper trade prospectively, then use a small live canary in a dedicated account. Compare intended orders with submitted orders and actual fills after every stage.

This framework tests whether an agent can operate a trading process, not whether one lucky run found alpha. If the goal is to build rather than select an agent, the same permission questions appear in the Claude and Alpaca MCP workflow.

What the test can—and cannot—prove

The Condor experiment proves that frontier AI agents can read accounts, use broker tools and route real orders under live conditions. It also shows that vague instructions can create large differences in turnover, risk behavior and tool use—and that broker capabilities can dominate a model comparison.

It does not prove that Claude is the best AI for trading, that Gemini is safest or that ChatGPT is unprofitable. The open-ended test had three observations per model, different broker pairings, tiny balances and no common benchmark or risk-adjusted evaluation. The agents are also non-deterministic, so a rerun may choose different trades.

The defensible conclusion is architectural: use an AI model for interpretation, research and planning; use deterministic code for permissions, sizing and risk; and use the broker only after both layers agree. Claude led this three-day snapshot, but the better trading system is the one that remains auditable and bounded when the model is wrong.

TAKE THE NEXT STEP

Build skills beyond the model leaderboard.

Learn to turn trading ideas into rules, code and backtests with Python and QuantConnect, while using AI as part of your learning workflow.