Understanding Results
Start here

Understanding Results

After an agent session completes, UCP Playground provides a detailed breakdown of what happened. Here is how to read each piece of data.

Outcome and funnel

Every session ends with an outcome label (how it ended) and a funnel reading (how far it got). The funnel is variant-aware — a retail session is scored against search → details → cart → checkout (a lodging session reads the same steps as search stays → hold room → book → booked), a procurement session against RFQ → quote → PO, and a session against a non-commerce server against a discovery-shaped ladder. The completion rate is the fraction of the resolved funnel's steps completed, and a confirmed purchase or booking counts as 100% even when middle steps were legitimately skipped.

A session counts as reaching checkout only when the store's response shows it: a checkout URL, or a checkout status or order returned by the store. A store-page (WebMCP) run is read against the same shopping funnel. It stops at the store's checkout and never pays, so it can't go further than checkout (see WebMCP: the store page).

The full outcome ladder and the funnel variants are documented in Funnel & Outcomes — this page stays at the reading level.

At a glance: outcomes at the top of the ladder (purchase_completed, checkout_reached, po_submitted) read as success; mid-funnel outcomes (cart_created, search_only, quote_received) read as partial progress; failed is reserved for actual errors. A cart_created session, for example, reads like this:

✓search returned actionable results8 items · ids + prices present
✓product details resolved a variantsize 10 · in stock
✓cart accepted the line itemline item + total present
·checkout completednot reached

Stages are ordered, and the list always renders complete. A · is not a pass — it means the stage was never reached, and the first gap explains everything below it.

Note

Some stores skip the details step entirely. If a store returns variant IDs directly in search results, a capable agent can jump from search straight to cart creation in fewer turns — the funnel records the steps that were actually completed.

Token Tracking

Each session records token consumption to help you estimate costs and compare model efficiency:

  • Prompt tokens -- Tokens sent to the model, including system prompts, tool definitions, and conversation history. Grows with each turn as context accumulates.
  • Completion tokens -- Tokens generated by the model for reasoning, tool calls, and responses. Lower is better for the same outcome.
  • Total tokens -- The sum of prompt and completion tokens. Use this to compare cost efficiency across models -- some models reach checkout in 10K tokens while others burn 60K.

Timing Metrics

Timing data shows how responsive the session was:

  • Duration -- Total wall-clock time from the first tool call to the final response
  • TTFB (Time to First Byte) -- How quickly the store's endpoint responded to the first request. High TTFB suggests network or server performance issues.
  • Turns -- The number of model invocations (request/response cycles). Fewer turns with the same outcome indicates a more efficient model.
Warning

Turn counts near the configured turn limit (8 by default) often signal that the agent is looping -- retrying failed tool calls or re-fetching the same data. Check the session timeline to diagnose what went wrong.

If the model provider cuts an answer off part-way, that turn is retried once; a second failure ends the run with an error instead of showing a half-finished answer.

Next Steps

If a session did not reach the outcome you expected, check the session timeline for specific errors. See Common Errors for guidance on the most frequent issues and how to resolve them.