o4-mini (Reasoning)
OpenAI's compact reasoning model. Currently non-functional for shopping — 92.9% tool error rate with almost every tool call failing. Appears to struggle with MCP tool schemas entirely. Not recommended until OpenAI addresses the tool-use reliability issues.
Avg Tokens70,680Avg Duration39.8sTurns to Checkout11.5
Shopping Score21/100
Weak
Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.
Shopping Score Breakdown
Checkout Rate
Cart Rate
Search Rate
Turn Efficiency
Token Efficiency
22% of sessions failed with errors.
Top Stores
| Store | Checkout % | Cart % |
|---|---|---|
| •••••••••••••••••••••••••••• | 66.7% | 66.7% |
| everlane.com | 60% | 60% |
| •••••••••••••••••••••••••••••••• | 33.3% | 33.3% |
| ••••••••••••••••••••••••••••• | 0% | 0% |
| craze.wtf | 0% | 0% |
| goswingofficial.com | 0% | 50% |
| ••••••••••••••••• | 0% | 0% |
| teveo.com | 0% | 50% |
| westwing.de | 0% | 33.3% |
| •••••••••••••••••••••••••••••••• | 0% | 100% |
Known Issues
Non-Functionalhigh
92.9% tool error rate — almost every tool call fails.
Schema Incompatibilityhigh
Cannot correctly format MCP tool arguments. Not usable for shopping agent tasks.
Token Usage
Avg per Session70,680
Fleet Average67,228
vs Fleet+5%
Median (p50)29,882
p90152,402
Prompt / Completion split
Prompt 67,948 (96%) Completion 2,732 (4%)
Daily Average Tokens (last 30 days)
Distribution range
Min0p2515,531p5029,882p7577,176p90152,402Max721,744