AnthropicAnthropic

Claude Opus 5

Anthropic's latest flagship (Claude Opus 5, succeeding the Opus 4.x line). Newly added — Playground benchmark metrics are still being gathered. The Opus line has historically posted the lowest tool error rate in our benchmark with a conservative, confirm-before-acting checkout style.

Avg Tokens134,187Avg Duration42.4sTurns to Checkout12
CheckoutCartSearchTurnsTokens1.823.654.500
Shopping Score11/100
Weak

Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.

Shopping Score Breakdown

Checkout Rate40%1.8%
Cart Rate20%23.6%
Search Rate10%54.5%
Turn Efficiency15%0%
Token Efficiency15%0%
45.5% of sessions failed with errors.

Top Stores

StoreCheckout %Cart %
•••••••••••••••••••••••25%100%
tattty.com0%100%
••••••••••••••••••••••••••0%33.3%
everlane.com0%0%
flower-shop.local0%0%
•••••••••••••••••••••••0%0%
••••••••••••••••••••••••••••••••0%0%

Known Issues

Metrics Pendinglow

Claude Opus 5 was recently added; its Playground benchmark metrics are still being gathered.

Token Usage

Avg per Session134,187
Fleet Average66,756
vs Fleet+101%
Median (p50)45,030
p90281,075
Prompt / Completion split
Prompt 132,536 (99%) Completion 1,652 (1%)
Daily Average Tokens (last 30 days)
Distribution range
Min0p2511,380p5045,030p75124,929p90281,075Max1,909,499

Test Claude Opus 5 on Your Store

Run a live agent session to see how Claude Opus 5 handles your store's checkout flow end-to-end.

Run Agent Session