ReferenceGlossary
One vocabulary across the product, the docs, and the API. When a term here conflicts with wording you see elsewhere, this page wins — tell us about the other place.
The four axes
Everything in the harness is one of four things: a Target (what you test), a Driver (who does the testing — a model plus an identity), a Check (what you assert), and a Result (what comes out).
Product & surfaces
| Term | Means |
|---|
| UCP Playground | The product — the whole test harness. |
| Inspector | The mode where you call tools by hand and read raw responses. |
| Agent | The mode where a model drives a real session against a target. |
| Shop All Stores | Beta Agent scope: instead of one store, the session connects to the multi-store resolver, the agent picks the store itself, and the session then hands off to that store to shop there directly. No domain needed. |
| Connection | The picker in the Agent page's endpoint card that sets how the agent reaches the store: UCP · MCP, UCP · REST, or, under Store page, WebMCP. Options the store doesn't offer are disabled. Switching between MCP and REST reconnects. |
| Store page | The store's own web page, as opposed to its UCP endpoint. A store-page run opens the page in a browser on our side and uses the tools the page registers. Signed-in users only; one task per run. |
| WebMCP | Tools a web page registers for agents through document.modelContext. Separate from UCP: UCP does not carry WebMCP, so it is never a UCP transport. The Inspector's WEBMCP view lists a page's tools with findings; the Agent page can run a task through them. See WebMCP: the store page. |
| API | The programmatic surface — the same harness, headless. |
| Buyer context | The spec's provisional buyer signals (country, region, postal code, language, currency) carried on catalog searches. Saved per account or workspace; hints for localisation, never identity. |
| Held state | The ledger of what merchants are holding for you -- carts, checkouts, holds, booking sessions -- one row per held item, mirrored from the merchant's own responses. Status and expiry are the merchant's; "Expired" means a call failed or a hold ran out. Release cancels a live booking session at the store; Forget and Dismiss only drop the row. |
| Booking session | The lodging counterpart of a checkout (dev.ucp.lodging.booking, draft): the rooms, dates and guests a buyer asked for, priced by the store, with when each amount is paid and the cancellation policy. It is held until it expires, then completed (booked) or canceled. |
| Negotiation | The connect-time readout of version selection and the capability intersection between a target's manifest and this platform, with a status and reason per capability. Snapshotted on every Session. |
| Runtime | Where a session's network legs execute: Cloud (our infrastructure — the default) or Local (your machine, via the local runner — coming). Web, API, CLI, and Desktop are surfaces that drive a runtime, not runtimes themselves. |
Spec alignment: Platform and Business
The UCP specification names the two parties as roles. We keep our product words, but here is the mapping so nothing needs translating between the spec and this harness:
| UCP term | Our term | Note |
|---|
| Platform | the harness itself | The intermediary initiating requests. The Playground is a Platform; its /.well-known/ucp-agent profile is a Platform profile. When you point your own agent at a reference store, you are the Platform. |
| Business | Target (Store) | The merchant receiving requests and advertising capabilities. Use "Business" when quoting the spec; "Target" in the product. |
| Payment Credential Provider | — | Tokenisation / payment processing. Outside the harness's boundary. |
| context (catalog search) | Buyer context | The spec's shopping/types/context.json signals. We send them where the target's tool accepts them and reflect whatever the target does with them. |
| Buyer / user | linked identity | The customer a Platform acts for — attached to a session via identity linking. |
Targets
| Term | Means |
|---|
| Target | Anything the harness points at. The noun for the thing under test — not "domain", not "merchant". |
| Store | A real commerce site in the world. A store becomes a target when you point the harness at it. |
| Reference store | A real UCP store we host, in a chosen enforcement posture. A demo target for store builders and the validation on-ramp for agent builders. Always pre-production. |
| Test store | A demo merchant the harness hosts where every tool works end to end, tagged in the try-pills: flower-shop.local for shopping and hotel.local for lodging. Test payment tokens only; nothing is charged. |
| Local target | A target on your own machine, reached over a tunnel today or the Local runtime when it lands. |
| Manifest | The /.well-known/ucp document — discovery and declared capabilities. |
| Environment | Whether a target is production or pre-production — a property of a target. Undeclared defaults to production, so the strictest probe gating applies. Distinct from the identity directory's production/sandbox split: a sandbox-registered key cannot verify against a production directory. |
| Transport | A way a store's manifest says it can be reached: MCP, REST, A2A, or embedded. The harness connects over MCP and REST. |
| A2A | Agent-to-Agent: a transport where the store exposes an agent rather than tools. The Inspector shows a declared A2A endpoint but connects over MCP and REST. |
| Embedded transport | A manifest entry with "transport": "embedded". It declares the store's Embedded Checkout (ECP), not a way to connect. |
| ECP | Embedded Checkout Protocol — a handoff, not a way to connect: how a store hands its checkout UI to a host application at the payment step. A capability observed at checkout. |
Drivers
| Term | Means |
|---|
| Model | The LLM (Claude, GPT, Gemini, …). Never "the agent" — the agent is the model plus the harness plus an identity. |
| Provider | The model's vendor. |
| Identity | The credential the agent presents to a store: Unsigned, UCP Platform Identity (the default), or Web Bot Auth (community — unsupported). The support labels are part of the name. |
Checks
| Term | Means |
|---|
| Conformance | The check package: schema quality + identity enforcement + funnel completion, measured live at runtime. A testing loop you fix against — never a saved report or certificate. |
| Schema quality | How well a target's tool schemas guide a model — seven weighted checks, graded A–F per tool and overall. |
| Identity enforcement | Whether a target actually enforces the identity scheme it declares — measured, not assumed from the declaration. |
| Funnel | How far a session got: search → details → cart → checkout (→ purchase on demo targets). A lodging session reads the same steps as search stays → hold room → book → booked. |
| Probe | One deliberately mutated request in a differential run. Active probes run only against reference stores, stores you own, or explicit opt-ins — never a cold third-party production endpoint. |
| Verdict | The output of any check. See below. |
Results
| Term | Means |
|---|
| Run | One execution (also the verb). |
| Session | The stored record of one run of one model against one target — every message, tool call, and event, replayable. |
| Replay | Playback of a session, shareable and embeddable. |
| Step screenshot | On a store-page (WebMCP) run, a picture of the page: when it opened, after each step that changed what a shopper sees (the cart after a cart change, a page a tool moved to), and at the checkout reached. Shown on the run and in its replays, with "What the checkout page showed" at the checkout. |
| Wire log | The Backend view of a replay — the machine-truth event log: leveled REQUEST/RESPONSE/ERROR rows with payloads and timing, as opposed to the Chat view's conversation. |
| Collection | A saved configuration of targets × models × sequences — the object you create, schedule, and re-run. Running one is an eval (category language, not an object). |
| Leaderboard | The public ranking of models — aggregated from every session run through the Playground: yours, other users', and automated collection runs alike. |
Verdict vocabulary
One language for every judgement; only the scale changes (grade, score, status, ratio). Underneath every check are five render states:
| State | Means |
|---|
pass | Measured, and it met the bar. |
partial | Measured, some criteria met. |
fail | Measured, the bar wasn't met. |
unknown | We could not measure reliably — never rendered as failure. |
not-run | Not attempted yet — distinct from both fail and unknown. |
Three rules apply everywhere: a pass state is a measured fact about the target (facts about our own request render neutral, never green); we report, never shame ("no gating observed", not "failed"); and unknown / not-run are always visually distinct from fail and from each other.
Identity enforcement verdicts
The identity check speaks a five-word domain vocabulary on top of those render states. Deliberately not pass/fail: a store that accepts everything is not failing a scheme, it is ignoring it.
| Verdict | Means | Renders as |
|---|
| Enforced | Valid requests accepted; every invalid variant treated differently. "Enforced" means the scheme has observable consequences — divergence, not necessarily a rejection. | pass |
| Partial | Some invalid variants are treated differently; others are accepted unchanged. | partial |
| Not enforced | No gating observed — signed and unsigned requests are treated identically. Only ever reported for a target that declares the scheme; a store that never claimed it has no identity verdict at all. | fail |
| Inconclusive | The valid signed request did not succeed, so enforcement could not be assessed. | unknown |
| Not run | Handshake only — the negative probes that produce a verdict have not been run. | not-run |
Postures
A posture is the stance a store is configured to take; a verdict is what the harness measures the behaviour to be. The interesting readouts are the disagreements — a store that believes it is Strict but measures Partial.
| Posture | Means |
|---|
| Strict | Valid → allow, missing → reject, invalid → reject. Recommended for production. |
| Permissive | Valid → allow, missing → allow, invalid → reject. The recommended rollout mode — an unsigned-allowed store is not a downgrade, and the missing-signature policy is reported as a characteristic, never folded into the verdict. |
| None | No enforcement. A reference-store posture for testing what "not enforced" looks like from the agent side. |