AlibabaAlibaba

QwQ 32B (Reasoning) Retired

This model is no longer available to run. Its stats below reflect past sessions and are kept for reference.

Alibaba's reasoning model. Reached the cart stage but did not complete checkouts in our benchmark — zero tool errors but also zero completions, with reasoning traces that over-analysed checkout requirements without acting. Retired: no longer served by the upstream router.

Avg Tokens14,586Avg Duration36.7s
CheckoutCartSearchTurnsTokens07.128.65050
Shopping Score19/100
Weak

Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.

Shopping Score Breakdown

Checkout Rate40%0%
Cart Rate20%7.1%
Search Rate10%28.6%
Turn Efficiency15%50%
Token Efficiency15%50%
71.4% of sessions failed with errors.

Top Stores

StoreCheckout %Cart %
allbirds.com0%0%
houseofparfum.nl0%33.3%
everlane.com0%0%
•••••••••••••••••••••••••••••0%0%
••••••••••••••••••••••••••0%50%
••••••••••••••••••••••••••••••••0%0%
••••••••••••••••••••••••••••••••0%0%

Known Issues

No known issues documented for this model yet.

Token Usage

Avg per Session14,586
Fleet Average67,666
vs Fleet−78%
Median (p50)0
p9026,283
Prompt / Completion split
Prompt 13,738 (94%) Completion 847 (6%)
No daily trend data yet. Run more sessions to see token usage over time.
Distribution range
Min0p250p500p7511,920p9026,283Max163,629

Test QwQ 32B (Reasoning) on Your Store

Run a live agent session to see how QwQ 32B (Reasoning) handles your store's checkout flow end-to-end.

Run Agent Session