Sessions & Replays
Test a store

Session History

Every agent session is saved automatically. Session History lets you browse, replay, compare, and share past sessions to track how stores and models perform over time.

Note

You must be logged in to access Session History. Sessions run while logged out are not persisted. Sign in to start saving your sessions.

Browsing Sessions

The Session History page lists all your past sessions, newest first. Each row shows key metrics at a glance:

  • Date -- when the session ran
  • Store -- the store that was tested (Shop All Stores sessions show as All Stores)
  • Model -- which AI model ran the session
  • Outcome -- the final result (purchase_completed, checkout_reached, cart_created, etc.). Lodging sessions show these as Booked, Booking Attempted and Room Held, and their steps as Search Stays, Hold Room, Book, Booked.
  • Steps -- a dot per funnel step, filled for each step the agent completed
  • Tokens -- total token usage (input + output) for the session
  • Duration -- wall-clock time from start to finish

Click anywhere on a row to open the session. Cmd- or Ctrl-click opens it in a new tab.

Filtering and Search

Use the filter bar to narrow down sessions by any combination of three criteria:

  • Store -- filter by domain to see all sessions for a specific store
  • Model -- isolate results for a particular AI model
  • Outcome -- show only sessions with a specific result (e.g., only failures to investigate issues)

Replaying Sessions

Open any session for the full replay view. A Chat / Backend toggle switches between two views of the same session:

  • Chat -- the conversation as it looked live: the agent's messages, its reasoning about which products to choose and why, and the products it surfaced.
  • Backend -- the Wire log: a dark pane of leveled event rows (REQUEST, RESPONSE, ERROR, THINKING, STEP, INFO), each with its tool name or label, payload, and a timestamp offset from session start. The header summarizes the rows beneath it -- "N events · N calls · N errors" -- and a Copy JSON button exports the raw steps. This is the machine-truth view of what actually went over the wire.

Above the conversation, a Connected with card records what the session started with: the Negotiation readout (version selection and the capability intersection, expandable to the full five steps), the Buyer context the searches carried, and the Held state this session produced -- carts, checkouts, or holds, read-only, as the merchant last reported them. These are snapshots taken at session start, so a replay shows what the session actually ran with, not what the merchant declares today. Sessions recorded before this card existed say so instead of showing empty blocks.

A store-page (WebMCP) run also shows what the browser saw: the page the agent opened and the checkout where it stopped at the top, and under each step that changed the page, a screenshot of the result (the cart after a cart change, or the page a tool moved to). At the checkout, What the checkout page showed opens the text of the store's checkout page. Screenshots are kept with the session. See WebMCP: the store page.

Your view choice sticks -- the toggle is remembered across replays. Replay plays the conversation back step by step. The replay is read-only: you are reviewing a completed session, not re-running it. Use the Backend view to verify a store's tool responses are well-formed, and the Chat view to understand why the agent made a decision.

Comparing Sessions

Tick the checkbox on two to five sessions in the history list, then click Compare Replay in the bar that appears. This places sessions side by side so you can directly compare how different models handled the same store, or how the same model performed before and after a store's UCP changes.

Each column's header also carries the session's negotiation line (v2026-04-08 · 7/7 active · shopping), so a difference in outcome can be read against a difference in what was negotiated. The comparison view highlights differences in funnel progress, token usage, number of turns, and outcome. This is especially useful for regression testing: run the same query after updating your store's UCP implementation and compare against the previous session.

Sharing Sessions

Any session can be turned into a shareable public link. Click the Share button on a session to generate a unique URL that anyone can view -- no login required. Shared sessions display the full replay in read-only mode, with personal details scrubbed. Sharing publishes a fixed copy of the session as it was when you shared it; to publish later turns of a session you continued, unshare it and share it again. The same copy is available as JSON at /s/{id}.json. The shared page keeps the negotiation readout and the buyer context reduced to country, currency, and language -- no region, no postal code -- and never includes your held-state ledger. A store-page run's screenshots are shown on the shared page too.

This is useful for:

  • Sending a session to a colleague to debug a store issue
  • Sharing results in a Slack thread or GitHub issue
  • Including evidence in a UCP compliance report

Embedding Sessions

For documentation, blog posts, or presentations, you can embed a session replay using an iframe. Once a session is shared, click the Embed button to copy a ready-to-use iframe snippet. The embedded replay is fully interactive -- viewers can scroll through the chat and inspect tool calls without leaving the page.

Tip

Combine sharing with the Funnel & Outcomes reading to create before-and-after case studies. Run a session, fix the issues it reveals, then run again and share both sessions to demonstrate the improvement.