Creating Collections
A collection defines what to test: which stores, which models, and what conversation sequences to run. Once saved, you can run it any time -- manually, via API, or on a schedule.
Choosing Stores
Add one or more store domains. Any store whose UCP manifest declares an MCP endpoint works -- Shopify stores, UCPReady (WooCommerce) stores, or custom UCP implementations. The eval runner discovers the MCP endpoint automatically from the store's /.well-known/ucp manifest.
Eval sessions run over UCP's MCP transport only. A store that declares REST but no MCP endpoint can't start an eval session; test it in the Agent or through the API instead. Runs through a store's own page tools (WebMCP) are not part of evals.
Start with stores you've already tested manually in the Agent. You know their tools work and what to expect. Then expand to new stores once your sequences are proven.
Choosing Models
Select which AI models to benchmark. All models from the Agent model list are available. Common patterns:
- Head-to-head: compare two models on the same stores (e.g., Gemini 3.5 Flash vs Gemini 3.1 Pro)
- Frontier sweep: test all frontier models to find the best for your use case
- Cost optimization: compare a frontier model against cheaper alternatives to find the best price/performance ratio
Writing Sequences
Sequences are the conversation scripts your eval runs. Each sequence has a name and a list of turns (user messages). Keep prompts generic so they work across different stores:
Good Sequences
- "What do you sell?" -- works on any store
- "Show me products under $60" -- works with any catalog
- "Add the first one to my cart" -- references the agent's previous response
- "Proceed to checkout" -- universal checkout trigger
Bad Sequences
- "Find the Nike Air Max 90 in size 10" -- too specific, only works on one store
- "Search, add to cart, checkout, and pay" -- too much in one turn, better as 3 separate turns
Session Count
The total number of sessions is: stores x models x sequences. Each session runs as an independent multi-turn conversation.
- 2 stores x 2 models x 1 sequence = 4 sessions
- 3 stores x 3 models x 2 sequences = 18 sessions
- 5 stores x 5 models x 1 sequence = 25 sessions
Each session consumes API tokens. A 25-session eval run uses roughly the same tokens as 25 individual agent sessions. Plan accordingly for large matrix runs.
API Example
Collections can be created via the API:
curl -X POST -H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
https://ucpplayground.com/api/v1/collections \
-d '{
"name": "Q1 Benchmark",
"config": {
"stores": ["oakywood.shop", "ugmonk.com"],
"models": ["gemini-3-flash", "gemini-3-1-pro"],
"sequences": [{
"name": "browse_and_buy",
"turns": [
{ "message": "Show me two products under $60" },
{ "message": "Add both to my cart" },
{ "message": "Proceed to checkout" }
]
}]
}
}'