QwQ 32B (Reasoning) Retired
This model is no longer available to run. Its stats below reflect past sessions and are kept for reference.
Alibaba's reasoning model. Reached the cart stage but did not complete checkouts in our benchmark — zero tool errors but also zero completions, with reasoning traces that over-analysed checkout requirements without acting. Retired: no longer served by the upstream router.
Avg Tokens14,586Avg Duration36.7s
Shopping Score19/100
Weak
Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.
Shopping Score Breakdown
Checkout Rate
Cart Rate
Search Rate
Turn Efficiency
Token Efficiency
71.4% of sessions failed with errors.
Top Stores
| Store | Checkout % | Cart % |
|---|---|---|
| allbirds.com | 0% | 0% |
| houseofparfum.nl | 0% | 33.3% |
| everlane.com | 0% | 0% |
| ••••••••••••••••••••••••••••• | 0% | 0% |
| •••••••••••••••••••••••••• | 0% | 50% |
| •••••••••••••••••••••••••••••••• | 0% | 0% |
| •••••••••••••••••••••••••••••••• | 0% | 0% |
Known Issues
No known issues documented for this model yet.
Token Usage
Avg per Session14,586
Fleet Average67,666
vs Fleet−78%
Median (p50)0
p9026,283
Prompt / Completion split
Prompt 13,738 (94%) Completion 847 (6%)
No daily trend data yet. Run more sessions to see token usage over time.
Distribution range
Min0p250p500p7511,920p9026,283Max163,629