DeepSeek R1 (Reasoning)
DeepSeek's reasoning model with chain-of-thought. Slow (76s avg) but deliberate — the reasoning traces show it carefully planning each tool call. Moderate error rate (12.5%) partly from JSON formatting issues. Small sample size but shows promise for complex checkout flows where thinking through validation errors matters.
Avg Tokens23,023Avg Duration53.7sTurns to Checkout6.7
Shopping Score23/100
Weak
Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.
Shopping Score Breakdown
Checkout Rate
Cart Rate
Search Rate
Turn Efficiency
Token Efficiency
61.5% of sessions failed with errors.
Top Stores
| Store | Checkout % | Cart % |
|---|---|---|
| allbirds.com | 33.3% | 33.3% |
| everlane.com | 33.3% | 33.3% |
| houseofparfum.nl | 20% | 40% |
| ••••••••••••••••••••••••••••• | 0% | 0% |
| •••••••••••••••••••••••••••••••• | 0% | 0% |
| ••••••••••••••••••••••••••••••• | 0% | 0% |
| •••••••••••••••••••••••••••••••• | 0% | 0% |
| •••••••••••••••••••••••••••••••• | 0% | 0% |
| ••••••••••••••••••••••••••••• | 0% | 0% |
| singon.com.np | 0% | 0% |
Known Issues
High Latencymedium
76 seconds average — reasoning overhead adds significant time.
JSON Formattingmedium
12.5% error rate partly from JSON formatting issues in tool arguments.
Token Usage
Avg per Session23,023
Fleet Average69,856
vs Fleet−67%
Median (p50)13,228
p9053,042
Prompt / Completion split
Prompt 21,831 (95%) Completion 1,192 (5%)
Daily Average Tokens (last 30 days)
Distribution range
Min2,618p257,311p5013,228p7532,490p9053,042Max116,012