Testing Your Store
A working manifest and endpoint are just the beginning. Thorough testing ensures AI agents can actually complete purchases on your store, not just connect to it.
Recommended Workflow
Follow this sequence to systematically validate your UCP integration from discovery through to checkout:
- Verify manifest discovery — Enter your domain in UCP Playground and confirm your
/.well-known/ucpmanifest is found and parsed correctly. Check that all service endpoints and payment handlers appear as expected. - Inspect each tool manually — Use the Tool Inspector to call each tool individually. Start with
search_catalog(orsearch_shop_catalogon older Shopify stores), thenget_product_details, then work through the cart and checkout flow. Verify that responses are valid JSON with the expected structure. - Fix schema issues — Look for common problems: missing required fields in your tool definitions, inconsistent ID formats between search results and detail lookups, prices in dollars versus cents, and missing variant IDs that agents need for cart operations.
- Run an agent session — Start a full agent session using a capable model like Claude Sonnet or GPT-5.5. Give it a realistic shopping task such as "Find a blue medium t-shirt and add it to the cart." Watch how the agent navigates your tools.
- Review the session timeline — After the session ends, examine the full timeline. Look at every tool call, the arguments the agent chose, and the responses your store returned. Pay attention to where the agent hesitated, retried, or gave up.
- Iterate — Fix any issues you find, retest the affected tools in the inspector, then run another agent session to verify the fix. Repeat until sessions consistently reach checkout.
- Check your conformance surface — Review your schema quality grade, and read Identity Enforcement to check who is calling your store and whether your identity posture behaves the way you intend.
Using the Tool Inspector
The tool inspector shows you exactly what AI agents see when they connect to your store. Every tool's input schema, description, and parameter constraints are visible. If your schema says a parameter is optional but your endpoint fails without it, you will catch that here before agents do.
Interpreting Session Outcomes
Session outcomes tell you where your integration breaks down. If agents consistently drop off at the same point, that points to a schema or implementation issue at that step:
- Agents stop after search — Your search results likely lack the IDs or URLs needed to fetch product details. Agents cannot proceed without a way to drill into specific products.
- Agents fail at cart creation — Check that your
create_carttool clearly documents what format variant IDs and quantities should be in. Mismatched ID formats between search results and cart inputs are the most common issue. - Agents stall at checkout — Checkout tools often require specific fields like buyer info, shipping addresses, or payment tokens. Make sure your tool descriptions explain what is required and in what format.
Testing Your Store Page (WebMCP)
If your storefront registers WebMCP tools on its own page (through document.modelContext), you can test those too. WebMCP is separate from UCP: it is the store page's own tools, not a UCP transport.
- In the Inspector — the WEBMCP button beside MCP and REST lists the tools your page registers, with findings on each, and a Page tools vs UCP tools table that matches them to your UCP tools by job. You can run read-only tools by hand.
- On the Agent page — signed in, choose WebMCP under Store page in the endpoint card's Connection picker. The agent runs one shopping task through your page's tools in a browser on our side. It stops at your checkout and never pays.
What it shows you: whether an agent can get from your page to your checkout using only the tools you register, and what a shopper would have seen on the way. The run keeps screenshots of the page it opened, of each step that changed what a shopper sees, and of the checkout it reached. Run the same task over UCP to see where the two routes differ. See WebMCP: the store page for limits and details.
Test Across Multiple Models
Different AI models interpret tool schemas differently. Run sessions with at least two or three models. If only one model fails at a particular step, the issue is likely model-specific and may resolve with prompt adjustments. If every model fails at the same point, the problem is in your store's schema or implementation and needs to be fixed on your end.
Share session replays with your engineering team using public share links. Each replay captures the full tool call history with request and response payloads, making it easy for developers to diagnose issues without needing to reproduce them.
For how stores are scored for UCP compliance, see UCP Checker, which grades stores and publishes results on its leaderboard.