Comparing Models
Test a store

Comparing Models

Four surfaces answer "which model?", at different altitudes: the leaderboard for aggregates, the Models Compare page for head-to-head stats, the Agent rail's Compare Models for live side-by-side runs, and session compare for finished sessions.

The four comparison surfaces

  • The Leaderboard -- aggregate rankings across every session run through the Playground. Best for the broad question: which models reliably reach checkout, and at what token cost.
  • Models Compare -- pick up to four models and see their aggregate stats side by side: checkout rate, tokens, duration, turns-to-checkout. Best for shortlisting before you run anything.
  • Compare Models in the Agent rail -- from a connected Agent session, the Compare Models button splits the chat into up to five models running the same store and query simultaneously. Best for watching how models diverge on your store, live.
  • Session compare -- select finished sessions from your history and place them side by side: funnel progress, tokens, turns, outcome. Best for before-and-after regression checks and post-hoc analysis.

Key Metrics to Compare

Whichever surface you use, focus on three dimensions:

  • Checkout Completion Rate -- The most important metric. A model with a 90% completion rate will successfully handle 9 out of 10 shopping requests. Even a few percentage points matter at scale.
  • Average Tokens -- Directly translates to cost. Token-efficient models reach checkout on a fraction of the tokens verbose models burn for the same task. At high volume, this difference is significant.
  • Average Duration -- How long users wait. Flash and mini variants are generally fastest; large reasoning models are slowest but can handle more complex catalogs. The leaderboard's duration column has the measured numbers.

Token Efficiency vs. Reliability

There is often a tradeoff between token efficiency and checkout completion rate. Smaller, faster models use fewer tokens per session but may fail on stores with complex product schemas -- for example, products with multiple option dimensions (color, size, material) that require careful variant resolution. Larger models tend to handle these edge cases better but at a higher token cost.

The leaderboard makes this tradeoff visible. Sort by completion rate to find the most reliable models, then check their token usage to understand the cost implications.

Store Compatibility

Not all models perform equally across all stores. A model that excels on Shopify stores might struggle with custom UCP implementations that return data in unexpected formats. The leaderboard aggregates across stores, so a model's overall score reflects its general versatility. However, if you are building for a specific store or platform, your own test sessions will be more informative than aggregate rankings.

Tip

Use the Agent rail's Compare Models with a specific, realistic prompt like "Find a blue cotton t-shirt in size medium under $40" rather than a vague one. Identical store, query, and timing gives you a controlled comparison on your exact use case.

Comparing routes, not models

On a store that also registers tools on its own page (WebMCP), you can hold the model fixed and change the route instead: run the same task over UCP and then through the store page, using the Agent page's Connection picker, and compare the two sessions. See WebMCP: the store page.

Building Your Own Benchmarks

The leaderboard is a starting point, not the final answer. Models are updated frequently, store implementations change, and your specific use case may weight metrics differently than the aggregate. Run sessions on the stores you care about with the prompts your users will actually send. For guidance on selecting models for different scenarios, see Choosing Models.