Running Sessions
An agent session sends an AI model into a real UCP-compatible store to shop autonomously. You connect a store, then chat -- the agent handles the tool calls.
Starting a Session
Open the Agent tab. The connect card asks for a model, a scope and a store; once connected you pick the connection, then chat:
- Model -- pick from the grouped-by-provider menu: 26 models across 8 providers (retired models stay on the leaderboard but can't start new runs). See Choosing Models for trade-offs.
- Scope -- a toggle between Specific Store and Shop All Stores (beta). Specific Store is the default.
- Store domain -- type the domain (e.g.,
everlane.com) and click Connect. Endpoint discovery verifies UCP support and fetches the store's manifest. A row of try-pills below offers known-good stores, withflower-shop.local(shopping) andhotel.local(lodging) tagged as test stores. - Connection -- once connected, the endpoint card shows how the agent reaches the store: UCP · MCP, UCP · REST, or, under Store page, WebMCP. A UCP option the store doesn't offer is disabled; WebMCP is listed when you are signed in. Switching between MCP and REST reconnects. The choice is fixed once the session starts (see Choosing the connection below).
- Chat -- type a shopping message in the composer and send it. There is no "Run" button; every message you send drives the agent. Be specific: "find me a blue t-shirt size M under $30" gives the agent criteria to act on.
What Happens Under the Hood
When a session starts, the agent connects to the store's UCP endpoint over the connection you chose. Over MCP, it uses JSON-RPC tool calls directly. Over REST (the default for REST-only stores), it generates tool definitions from the manifest capabilities and routes calls through the REST transport. The agent decides which tools to call, in what order, and with what parameters -- there is no human-in-the-loop during execution.
A typical session follows the funnel: search the catalog, view product details, add items to a cart, and attempt checkout. Against a lodging store the same run searches stays, holds a room in a booking session and books it, and the chat shows hotel cards for the rooms and the booking. The agent tracks its own progress and adapts based on tool responses. Text written by the merchant or third parties inside tool replies (descriptions, reviews, store notices) is treated as information, not instructions; a tool's own guidance on which tool to call next is followed. If the model provider cuts an answer off part-way, that turn is retried once. Sessions save automatically when the agent finishes or reaches a terminal state -- review any past session in Sessions & Replays.
The rail
The info rail beside the chat carries the session's state, top to bottom:
- Endpoint status -- the resolved endpoint URL, its status badge, the last ping, and the Connection picker. Once connected, a Save as target button below it stores the domain and discovered endpoints as a reusable target.
- Runtime -- where merchant-bound requests execute: Cloud (default) or Local via a connected runner. See Localhost & Tunnels.
- Agent Identity -- the signing identity the session presents to the store (when enabled for your workspace). Stores that support identity linking also surface a Link Identity button here.
- Buyer context -- the spec
contextsignals (country, currency, language, optional region and postal code) the agent's catalog searches carry. Click the pill to change them inline; the same pill sits beside the composer. Whether they change prices or availability is the merchant's call -- the tool result shows what came back. See Buyer context, held state, and negotiation. - Held state -- everything a merchant has issued an id for on your behalf: carts, checkouts, holds, booking sessions. One row each, mirrored from the merchant's own response: the store, its total, and one status (Cart, Checkout, Held, Ready, Booked, Needs buyer, or Expired) with a live countdown when the merchant gave an expiry. Open a row for its id, the tool that last synced it, the merchant's message codes (such as
local_tax) and its line items. Expired rows are grouped and collapsed, with Clear. Release cancels a live booking session at the store; Forget and Dismiss only drop the row here and never contact the merchant. - Negotiation -- how the platform negotiated with the store's manifest: version selection, the capability intersection, the resulting prompt variant, and a status with a reason on every capability row. Collapsed by default; the badge reads
active/declared · variant. - Schema Quality -- the seven-check grader's verdict on the store's tool schemas: an overall letter grade plus a per-tool letter grade and issue list. See Schema Quality.
- Funnel Progress -- the resolved funnel's steps, checked off as the agent completes them.
- Session Stats -- tokens (with a meter against the run cap), turns, and the outcome once one is classified.
- Model -- the same grouped-by-provider menu, shown until the session starts, with a Compare Models button beneath it.
Choosing the connection
- UCP · MCP and UCP · REST -- the store's UCP endpoint, as declared in its manifest. These are conversations: ask follow-ups in the same session.
- Store page · WebMCP -- the tools the store's own web page registers, run in a browser on our side. This is not a UCP transport. It needs you to be signed in, runs one task per session (click New run for the next), allows 10 store-page runs per day, stops at the store's checkout and never pays. The run shows the page it opened and the checkout it reached, and its replay keeps a screenshot per step. See WebMCP: the store page.
Shop All Stores (beta)
Flip the Scope toggle to Shop All Stores and no domain is needed. The session runs in two stages: it connects to the multi-store resolver — a read-only index of verified UCP stores on every platform, drawn from UCP Checker — and the agent searches the catalogue, checks each candidate's observed state and policies, and picks a store. It then hands off to that store, connecting to the merchant's own endpoint and continuing the same session with the store's own tools, including its cart and checkout. The purchase runs Playground to merchant directly; the index never carries it. If the store it picks can't be reached, the agent keeps the index tools and can choose another. Results vary while this mode is in beta.
Running Multiple Models
The Compare Models button in the rail splits the session into up to 5 models simultaneously against the same store and query. Each model runs its own independent session with its own tool calls, so results are directly comparable -- watch them progress side by side through the funnel.
Keep your queries specific for the best results. "Find me Nike Air Max size 10 in white" gives the agent clear criteria to search and filter on. Vague queries like "show me shoes" lead to ambiguous decisions and lower completion rates.
If a session stalls or takes longer than expected, check the timeline. Common causes include the store returning unexpected response formats or the agent struggling with variant selection. See Troubleshooting for help.