xAIxAI

Grok 4.5

xAI's latest flagship (Grok 4.5, xAI's most capable model). A 500K-context model positioned for coding, knowledge work, and STEM, priced above the rest of the Grok 4.x line. Newly added — Playground benchmark metrics are still being gathered. The Grok line has historically posted a standout reliability record in our benchmark (near-zero session failures, low tool error rate) with latency as its main weakness.

Avg Tokens105,103Avg Duration27.5sTurns to Checkout8
CheckoutCartSearchTurnsTokens4.540.990.95082
Shopping Score39/100
Fair

Scores computed from real agent sessions against live UCP-enabled stores. Not estimated — every data point is from an actual tool call, checkout attempt, and store response.

Shopping Score Breakdown

Checkout Rate40%4.5%
Cart Rate20%40.9%
Search Rate10%90.9%
Turn Efficiency15%50%
Token Efficiency15%82%
9.1% of sessions failed with errors.

Top Stores

StoreCheckout %Cart %
teveo.com0%100%
••••••••••••••••••••••••••••••••0%0%
yourkaya.pl0%50%

Known Issues

Metrics Pendinglow

Grok 4.5 was recently added; its Playground benchmark metrics are still being gathered. The Grok line has historically shown high latency as its main weakness.

Token Usage

Avg per Session105,103
Fleet Average66,756
vs Fleet+57%
Median (p50)50,357
p90306,511
Prompt / Completion split
Prompt 104,132 (99%) Completion 972 (1%)
Daily Average Tokens (last 30 days)
Distribution range
Min5,610p2533,858p5050,357p75113,628p90306,511Max375,414

Test Grok 4.5 on Your Store

Run a live agent session to see how Grok 4.5 handles your store's checkout flow end-to-end.

Run Agent Session