Claude Sonnet 4.6
Anthropic's prior-generation balanced model (Sonnet 4.6, superseded by Sonnet 5), kept as a lower-cost option. Strong multi-turn checkout flow with wallet payment and identity linking support; higher token usage than Opus due to verbose tool responses.
Avg Tokens70,823Avg Duration34.1sTurns to Checkout8.4
Shopping Score39/100
Fair
Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.
Shopping Score Breakdown
Checkout Rate
Cart Rate
Search Rate
Turn Efficiency
Token Efficiency
19.5% of sessions failed with errors.
Top Stores
| Store | Checkout % | Cart % |
|---|---|---|
| •••••••••••••••• | 100% | 100% |
| houseofparfum.nl | 75.5% | 85.9% |
| forever21.com | 66.7% | 66.7% |
| ••••••••••••••••••• | 62.5% | 62.5% |
| •••••••••••••••••••••••••••••••• | 54.5% | 54.5% |
| •••••••••••••••••••••••••••••••• | 50% | 100% |
| mr-fothergills.co.uk | 50% | 50% |
| westwing.de | 50% | 50% |
| reebok.com | 50% | 50% |
| •••••••••••••••••••••••••••••••• | 50% | 100% |
Known Issues
Nudge Dependencymedium
Stops after validation errors 9.3% of the time. Needs orchestrator nudge to continue checkout.
High Token Usagelow
Averages 67K tokens per session — highest among Anthropic models. Verbose tool response handling.
Meta Stringificationmedium
Frequently stringifies meta objects instead of passing as JSON. Requires orchestrator auto-injection for idempotency keys.
Token Usage
Avg per Session70,823
Fleet Average66,593
vs Fleet+6%
Median (p50)39,866
p90133,329
Prompt / Completion split
Prompt 69,459 (98%) Completion 1,065 (2%)
Daily Average Tokens (last 30 days)
Distribution range
Min0p2519,558p5039,866p7577,502p90133,329Max1,935,784