AnthropicAnthropic

Claude Sonnet 5

Anthropic's latest balanced model (Claude Sonnet 5, succeeding the Sonnet 4.x line). Newly added — Playground benchmark metrics are still being gathered. The Sonnet line historically pairs a strong multi-turn checkout flow with faster, more token-efficient responses than Opus.

Avg Tokens386,011Avg Duration69.8sTurns to Checkout33.2
CheckoutCartSearchTurnsTokens10388600
Shopping Score20/100
Weak

Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.

Shopping Score Breakdown

Checkout Rate40%10%
Cart Rate20%38%
Search Rate10%86%
Turn Efficiency15%0%
Token Efficiency15%0%
14% of sessions failed with errors.

Top Stores

StoreCheckout %Cart %
•••••••••••••••••••••••••••••80%100%
••••••••••••••••••••20%20%
••••••••••••••••••••••••••••••••0%0%
•••••••••••••••••0%0%
••••••••••••••••••••••••••0%50%
flower-shop.local0%33.3%
allbirds.com0%50%
houseofparfum.nl0%0%
••••••••••••••••••••••••••••••••0%50%
terryshomegoods.com0%100%

Known Issues

Metrics Pendinglow

Claude Sonnet 5 was recently added; its Playground benchmark metrics are still being gathered.

Token Usage

Avg per Session386,011
Fleet Average66,756
vs Fleet+478%
Median (p50)95,511
p90903,108
Prompt / Completion split
Prompt 383,527 (99%) Completion 2,484 (1%)
Daily Average Tokens (last 30 days)
Distribution range
Min0p2524,992p5095,511p75336,312p90903,108Max4,631,643

Test Claude Sonnet 5 on Your Store

Run a live agent session to see how Claude Sonnet 5 handles your store's checkout flow end-to-end.

Run Agent Session