GPT-4o
OpenAI's workhorse model. Fast (15.6s avg) and token-efficient but lower checkout completion than frontier models. Moderate tool error rate (8.4%). Never needs nudging — either completes or gives up cleanly. Good for high-volume, cost-sensitive benchmarking where speed matters more than checkout depth.
Avg Tokens34,077Avg Duration16.4sTurns to Checkout8.2
Shopping Score32/100
Weak
Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.
Shopping Score Breakdown
Checkout Rate
Cart Rate
Search Rate
Turn Efficiency
Token Efficiency
26.6% of sessions failed with errors.
Top Stores
| Store | Checkout % | Cart % |
|---|---|---|
| everlane.com | 61.5% | 61.5% |
| allbirds.com | 33.3% | 33.3% |
| ••••••••••••••• | 33.3% | 33.3% |
| houseofparfum.nl | 30.8% | 38.5% |
| ucp.travel | 28.6% | 28.6% |
| kyliecosmetics.com | 20% | 20% |
| •••••••••••••••••••••••••••••••• | 9.1% | 45.5% |
| ••••••••••••••••••• | 8.3% | 91.7% |
| •••••••••••••••••••••••••• | 0% | 50% |
| •••••••••••••••••••••••••••••••• | 0% | 0% |
Known Issues
Low Checkout Completionmedium
33% checkout rate despite fast execution. Gives up rather than retrying on errors.
Moderate Error Ratemedium
8.4% tool error rate — higher than Claude and Gemini models.
Token Usage
Avg per Session34,077
Fleet Average68,722
vs Fleet−50%
Median (p50)22,074
p9085,285
Prompt / Completion split
Prompt 33,535 (98%) Completion 542 (2%)
Daily Average Tokens (last 30 days)
Distribution range
Min0p258,169p5022,074p7539,124p9085,285Max437,563