AnthropicAnthropic

Claude Opus 4.8

Anthropic's previous flagship (Opus 4.8, superseded by Opus 5), kept as a lower-cost option. The Opus line has the lowest tool error rate in our benchmark (~2%) with efficient checkout completion (~7 turns) and a conservative, confirm-before-acting style. The model family that completed the first fully autonomous wallet-funded purchase through UCP.

Avg Tokens61,598Avg Duration30.7sTurns to Checkout8.2
CheckoutCartSearchTurnsTokens23.331.562.9482
Shopping Score29/100
Weak

Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.

Shopping Score Breakdown

Checkout Rate40%23.3%
Cart Rate20%31.5%
Search Rate10%62.9%
Turn Efficiency15%48%
Token Efficiency15%2%
36.8% of sessions failed with errors.

Top Stores

StoreCheckout %Cart %
forever21.com100%100%
••••••••••••••••••••••••••••100%100%
sini.fi100%100%
••••••••••••••••••••••••••••••••70%90%
••••••••••••••••••••••••••••••••66.7%66.7%
•••••••••••••••••••62.5%62.5%
ucp.travel61.5%61.5%
houseofparfum.nl58.5%69.8%
jbhifi.com.au50%50%
thehouseofrare.com50%50%

Known Issues

Nudge Dependencylow

Occasionally needs nudge (6.8%) but recovers cleanly and completes checkout.

Conservative Behaviorlow

Prefers to ask for confirmation rather than proceeding autonomously. Adds turns but reduces errors.

Token Usage

Avg per Session61,598
Fleet Average67,346
vs Fleet−9%
Median (p50)27,477
p90161,344
Prompt / Completion split
Prompt 60,295 (98%) Completion 1,015 (2%)
Daily Average Tokens (last 30 days)
Distribution range
Min0p2510,695p5027,477p7571,020p90161,344Max928,724

Test Claude Opus 4.8 on Your Store

Run a live agent session to see how Claude Opus 4.8 handles your store's checkout flow end-to-end.

Run Agent Session