Grok 4.3 Retired
This model is no longer available to run. Its stats below reflect past sessions and are kept for reference.
xAI's Grok 4.3 (this slot also carries earlier Grok 4 runs). Standout reliability in our benchmark — near-zero session failures and a very low tool error rate — with an efficient checkout path; latency was the main drawback, the slowest of the frontier models. Retired: superseded by Grok 4.5.
Avg Tokens35,601Avg Duration45.4sTurns to Checkout6.5
Shopping Score43/100
Fair
Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.
Shopping Score Breakdown
Checkout Rate
Cart Rate
Search Rate
Turn Efficiency
Token Efficiency
35.6% of sessions failed with errors.
Top Stores
| Store | Checkout % | Cart % |
|---|---|---|
| ••••••••••••••••••• | 100% | 100% |
| ucp.travel | 100% | 100% |
| ••••••••••••••• | 80% | 100% |
| •••••••••••••••••••••••••••••••• | 50% | 100% |
| houseofparfum.nl | 46.4% | 50% |
| everlane.com | 33.3% | 33.3% |
| allbirds.com | 0% | 0% |
| ••••••••••••••••••••••••••••• | 0% | 0% |
| ••••••••••••••••••••••••••••• | 0% | 0% |
| •••••••••••••••••••••••••••••••• | 0% | 0% |
Known Issues
High Latencymedium
84 seconds average per session — slowest frontier model by a significant margin.
Token Usage
Avg per Session35,601
Fleet Average66,570
vs Fleet−47%
Median (p50)22,188
p9093,360
Prompt / Completion split
Prompt 33,765 (95%) Completion 1,836 (5%)
No daily trend data yet. Run more sessions to see token usage over time.
Distribution range
Min0p259,814p5022,188p7542,505p9093,360Max401,007