Glossary
Reference

Glossary

One vocabulary across the product, the docs, and the API. When a term here conflicts with wording you see elsewhere, this page wins — tell us about the other place.

The four axes

Everything in the harness is one of four things: a Target (what you test), a Driver (who does the testing — a model plus an identity), a Check (what you assert), and a Result (what comes out).

Product & surfaces

TermMeans
UCP PlaygroundThe product — the whole test harness.
InspectorThe mode where you call tools by hand and read raw responses.
AgentThe mode where a model drives a real session against a target.
Shop All StoresBeta Agent scope: instead of one store, the session connects to the multi-store resolver, the agent picks the store itself, and the session then hands off to that store to shop there directly. No domain needed.
ConnectionThe picker in the Agent page's endpoint card that sets how the agent reaches the store: UCP · MCP, UCP · REST, or, under Store page, WebMCP. Options the store doesn't offer are disabled. Switching between MCP and REST reconnects.
Store pageThe store's own web page, as opposed to its UCP endpoint. A store-page run opens the page in a browser on our side and uses the tools the page registers. Signed-in users only; one task per run.
WebMCPTools a web page registers for agents through document.modelContext. Separate from UCP: UCP does not carry WebMCP, so it is never a UCP transport. The Inspector's WEBMCP view lists a page's tools with findings; the Agent page can run a task through them. See WebMCP: the store page.
APIThe programmatic surface — the same harness, headless.
Buyer contextThe spec's provisional buyer signals (country, region, postal code, language, currency) carried on catalog searches. Saved per account or workspace; hints for localisation, never identity.
Held stateThe ledger of what merchants are holding for you -- carts, checkouts, holds, booking sessions -- one row per held item, mirrored from the merchant's own responses. Status and expiry are the merchant's; "Expired" means a call failed or a hold ran out. Release cancels a live booking session at the store; Forget and Dismiss only drop the row.
Booking sessionThe lodging counterpart of a checkout (dev.ucp.lodging.booking, draft): the rooms, dates and guests a buyer asked for, priced by the store, with when each amount is paid and the cancellation policy. It is held until it expires, then completed (booked) or canceled.
NegotiationThe connect-time readout of version selection and the capability intersection between a target's manifest and this platform, with a status and reason per capability. Snapshotted on every Session.
RuntimeWhere a session's network legs execute: Cloud (our infrastructure — the default) or Local (your machine, via the local runner — coming). Web, API, CLI, and Desktop are surfaces that drive a runtime, not runtimes themselves.

Spec alignment: Platform and Business

The UCP specification names the two parties as roles. We keep our product words, but here is the mapping so nothing needs translating between the spec and this harness:

UCP termOur termNote
Platformthe harness itselfThe intermediary initiating requests. The Playground is a Platform; its /.well-known/ucp-agent profile is a Platform profile. When you point your own agent at a reference store, you are the Platform.
BusinessTarget (Store)The merchant receiving requests and advertising capabilities. Use "Business" when quoting the spec; "Target" in the product.
Payment Credential Provider—Tokenisation / payment processing. Outside the harness's boundary.
context (catalog search)Buyer contextThe spec's shopping/types/context.json signals. We send them where the target's tool accepts them and reflect whatever the target does with them.
Buyer / userlinked identityThe customer a Platform acts for — attached to a session via identity linking.

Targets

TermMeans
TargetAnything the harness points at. The noun for the thing under test — not "domain", not "merchant".
StoreA real commerce site in the world. A store becomes a target when you point the harness at it.
Reference storeA real UCP store we host, in a chosen enforcement posture. A demo target for store builders and the validation on-ramp for agent builders. Always pre-production.
Test storeA demo merchant the harness hosts where every tool works end to end, tagged in the try-pills: flower-shop.local for shopping and hotel.local for lodging. Test payment tokens only; nothing is charged.
Local targetA target on your own machine, reached over a tunnel today or the Local runtime when it lands.
ManifestThe /.well-known/ucp document — discovery and declared capabilities.
EnvironmentWhether a target is production or pre-production — a property of a target. Undeclared defaults to production, so the strictest probe gating applies. Distinct from the identity directory's production/sandbox split: a sandbox-registered key cannot verify against a production directory.
TransportA way a store's manifest says it can be reached: MCP, REST, A2A, or embedded. The harness connects over MCP and REST.
A2AAgent-to-Agent: a transport where the store exposes an agent rather than tools. The Inspector shows a declared A2A endpoint but connects over MCP and REST.
Embedded transportA manifest entry with "transport": "embedded". It declares the store's Embedded Checkout (ECP), not a way to connect.
ECPEmbedded Checkout Protocol — a handoff, not a way to connect: how a store hands its checkout UI to a host application at the payment step. A capability observed at checkout.

Drivers

TermMeans
ModelThe LLM (Claude, GPT, Gemini, …). Never "the agent" — the agent is the model plus the harness plus an identity.
ProviderThe model's vendor.
IdentityThe credential the agent presents to a store: Unsigned, UCP Platform Identity (the default), or Web Bot Auth (community — unsupported). The support labels are part of the name.

Checks

TermMeans
ConformanceThe check package: schema quality + identity enforcement + funnel completion, measured live at runtime. A testing loop you fix against — never a saved report or certificate.
Schema qualityHow well a target's tool schemas guide a model — seven weighted checks, graded A–F per tool and overall.
Identity enforcementWhether a target actually enforces the identity scheme it declares — measured, not assumed from the declaration.
FunnelHow far a session got: search → details → cart → checkout (→ purchase on demo targets). A lodging session reads the same steps as search stays → hold room → book → booked.
ProbeOne deliberately mutated request in a differential run. Active probes run only against reference stores, stores you own, or explicit opt-ins — never a cold third-party production endpoint.
VerdictThe output of any check. See below.

Results

TermMeans
RunOne execution (also the verb).
SessionThe stored record of one run of one model against one target — every message, tool call, and event, replayable.
ReplayPlayback of a session, shareable and embeddable.
Step screenshotOn a store-page (WebMCP) run, a picture of the page: when it opened, after each step that changed what a shopper sees (the cart after a cart change, a page a tool moved to), and at the checkout reached. Shown on the run and in its replays, with "What the checkout page showed" at the checkout.
Wire logThe Backend view of a replay — the machine-truth event log: leveled REQUEST/RESPONSE/ERROR rows with payloads and timing, as opposed to the Chat view's conversation.
CollectionA saved configuration of targets × models × sequences — the object you create, schedule, and re-run. Running one is an eval (category language, not an object).
LeaderboardThe public ranking of models — aggregated from every session run through the Playground: yours, other users', and automated collection runs alike.

Verdict vocabulary

One language for every judgement; only the scale changes (grade, score, status, ratio). Underneath every check are five render states:

StateMeans
passMeasured, and it met the bar.
partialMeasured, some criteria met.
failMeasured, the bar wasn't met.
unknownWe could not measure reliably — never rendered as failure.
not-runNot attempted yet — distinct from both fail and unknown.

Three rules apply everywhere: a pass state is a measured fact about the target (facts about our own request render neutral, never green); we report, never shame ("no gating observed", not "failed"); and unknown / not-run are always visually distinct from fail and from each other.

Identity enforcement verdicts

The identity check speaks a five-word domain vocabulary on top of those render states. Deliberately not pass/fail: a store that accepts everything is not failing a scheme, it is ignoring it.

VerdictMeansRenders as
EnforcedValid requests accepted; every invalid variant treated differently. "Enforced" means the scheme has observable consequences — divergence, not necessarily a rejection.pass
PartialSome invalid variants are treated differently; others are accepted unchanged.partial
Not enforcedNo gating observed — signed and unsigned requests are treated identically. Only ever reported for a target that declares the scheme; a store that never claimed it has no identity verdict at all.fail
InconclusiveThe valid signed request did not succeed, so enforcement could not be assessed.unknown
Not runHandshake only — the negative probes that produce a verdict have not been run.not-run

Postures

A posture is the stance a store is configured to take; a verdict is what the harness measures the behaviour to be. The interesting readouts are the disagreements — a store that believes it is Strict but measures Partial.

PostureMeans
StrictValid → allow, missing → reject, invalid → reject. Recommended for production.
PermissiveValid → allow, missing → allow, invalid → reject. The recommended rollout mode — an unsigned-allowed store is not a downgrade, and the missing-signature policy is reported as a characteristic, never folded into the verdict.
NoneNo enforcement. A reference-store posture for testing what "not enforced" looks like from the agent side.