GPT-5.5
OpenAI's frontier model (GPT-5.5, succeeding 5.2). The prior 5.2 revision had a notably high tool-error rate in our benchmark — nearly half of tool calls failed against MCP schemas, though turn efficiency was reasonable (~7.6 turns) when it reached checkout. 5.5 is the newer revision and is being re-benchmarked.
Avg Tokens52,623Avg Duration33.0sTurns to Checkout10.1
Shopping Score21/100
Weak
Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.
Shopping Score Breakdown
Checkout Rate
Cart Rate
Search Rate
Turn Efficiency
Token Efficiency
31.3% of sessions failed with errors.
Top Stores
| Store | Checkout % | Cart % |
|---|---|---|
| •••••••••••••••••••••••••••••••• | 100% | 100% |
| •••••••••••••••••••••••••••••••• | 85.7% | 85.7% |
| •••••••••••••••••••••••••••••••• | 50% | 50% |
| kyliecosmetics.com | 50% | 50% |
| ••••••••••••••••••• | 37.5% | 37.5% |
| •••••••••••••••••••••••••••••••• | 33.3% | 33.3% |
| houseofparfum.nl | 30.4% | 39.1% |
| apneeswimwear.com | 25% | 25% |
| everlane.com | 22.2% | 22.2% |
| ucp.travel | 20% | 20% |
Known Issues
High Tool Error Ratehigh
46.9% of all tool calls fail. Nearly half of MCP interactions produce errors.
Schema Compliancehigh
Struggles with MCP tool schemas — incorrect argument types, missing required fields.
Token Usage
Avg per Session52,623
Fleet Average66,736
vs Fleet−21%
Median (p50)24,426
p90129,545
Prompt / Completion split
Prompt 51,512 (98%) Completion 1,111 (2%)
Daily Average Tokens (last 30 days)
Distribution range
Min0p259,230p5024,426p7555,282p90129,545Max601,707