How Scoring Works
Conformance

How Scoring Works

The leaderboard ranks AI models by how well they complete real shopping tasks -- not synthetic benchmarks, but actual sessions against live stores.

Primary Metric: Checkout Completion Rate

Models are ranked primarily by their checkout completion rate -- the percentage of agent sessions whose outcome is checkout_reached or purchase_completed (a confirmed purchase or booking counts, naturally, as having reached checkout). A model that gets there in 8 out of 10 sessions scores 80%. This single metric captures the full picture: a model must search effectively, select the right product, handle variants, build a cart, and navigate the checkout flow without getting stuck or looping.

Completion rate is the clearest signal of real-world usefulness. A model that writes eloquent product descriptions but cannot add an item to a cart is not helpful for commerce tasks.

Secondary Metrics

Beyond completion rate, three secondary metrics provide additional context for comparison:

  • Average Tokens Consumed -- How many tokens the model uses per session. Lower token usage means lower cost per transaction. Some models reach checkout in 10K tokens while others need 60K for the same task.
  • Average Session Duration -- Wall-clock time from the first tool call to the final response. This reflects both model inference speed and how many round trips the model needs.
  • Average Turns -- The number of tool call rounds in a session. Fewer turns generally means the model is making better decisions about which tools to call and in what order.

Where the Data Comes From

The leaderboard is aggregated from every session run through the Playground -- yours, other users', anonymous runs, and automated collection runs alike. When anyone runs a session -- choosing a model, entering a shopping prompt, and letting the agent work against a live store -- that session contributes to the leaderboard data. There is no separate curated benchmark set, and no simulated responses or pre-scripted interactions. Runs through a store's own page tools (WebMCP) are the one exception: they are not part of the leaderboard, and their results stay with the person who ran them.

This means scores reflect genuine model behavior: handling unexpected product schemas, recovering from tool errors, and dealing with the variance that comes from real store data. Scores update continuously as new sessions complete, so the leaderboard always reflects the latest model performance.

Note

Scores can shift as stores update their UCP implementations or as model providers release new versions. A model's ranking today may differ from last week if the underlying model has been updated.

Understanding the Funnel

The checkout completion rate measures the final stage of the shopping funnel. To understand how sessions break down across each stage -- from search through cart creation to checkout -- see the Funnel & Outcomes guide.