DeepSeek V3.2
The highest checkout rate in the benchmark at 68%. Efficient at 6.2 turns average with low token usage. The trade-off: a 13.1% tool error rate, suggesting it brute-forces through errors rather than avoiding them. Never needs nudging — pushes through autonomously. The best raw completion rate of any model tested.
Avg Tokens69,072Avg Duration47.5sTurns to Checkout6.7
Shopping Score44/100
Fair
Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.
Shopping Score Breakdown
Checkout Rate
Cart Rate
Search Rate
Turn Efficiency
Token Efficiency
22.9% of sessions failed with errors.
Top Stores
| Store | Checkout % | Cart % |
|---|---|---|
| ••••••••••••••• | 77.8% | 77.8% |
| houseofparfum.nl | 72% | 72% |
| monos.com | 50% | 50% |
| ucp.travel | 33.3% | 33.3% |
| allbirds.com | 20% | 20% |
| snocks.com | 0% | 0% |
| ••••••••••• | 0% | 50% |
| ••••••••••••••••• | 0% | 0% |
| •••••••••••••••••••••••••••••••• | 0% | 100% |
| •••••••••••••••• | 0% | 50% |
Known Issues
Tool Error Ratemedium
13.1% error rate — brute-forces through errors rather than avoiding them.
JSON Formattinglow
Occasional issues with tool argument types. Requires extra validation overlays.
Token Usage
Avg per Session69,072
Fleet Average67,345
vs Fleet+3%
Median (p50)31,876
p90148,656
Prompt / Completion split
Prompt 68,312 (99%) Completion 760 (1%)
Daily Average Tokens (last 30 days)
Distribution range
Min2,517p2518,589p5031,876p7559,990p90148,656Max698,476