Reading Reports
Every eval run produces a structured report. Here's how to read the funnel matrix, performance table, and PDF benchmark report.
Funnel Comparison Matrix
The funnel matrix shows whether each store+model combination passed or failed each shopping stage:
- Search -- did the agent successfully call the search tool and get results?
- Details -- did the agent fetch product details? (may show FAIL if skipped but not needed)
- Cart -- did the agent add items to a cart?
- Checkout -- did the session reach checkout? It counts only when the store's response shows it: a checkout URL, or a checkout status or order returned by the store.
A "Details: FAIL" in the funnel doesn't always mean failure. Some stores include full variant data in search results, so the agent skips the details step and goes straight to cart. If Cart and Checkout both pass, the Details skip is fine.
Performance Table
For each session, the report shows:
- Tokens -- total tokens consumed (prompt + completion). Lower is cheaper.
- Duration -- wall-clock time for the full multi-turn session. Includes model inference + tool call latency.
- Turns -- total orchestrator turns across all sequence steps. Fewer turns = more efficient model.
- Cart Value -- total cart value if the agent reached cart creation. Shows currency.
Recommendations
The report auto-generates actionable recommendations based on the results:
- Store-level: "example-store.com: checkout rate is 50%. Review MCP tool responses for missing checkout URLs."
- Token usage: "Average token usage is 95K -- above the 40K baseline. Consider truncating product descriptions in MCP responses."
- Duration: "Average session duration is 48s -- above 15s target. Optimise MCP endpoint response times."
When all sessions reach checkout with zero errors, the report says: "All stores are fully AI-agent ready."
PDF Report
Every run can be downloaded as a PDF -- a 2-page "AI Commerce Benchmark Report" containing:
- Header with date, UCP version, models, stores, and session count
- Headline metrics: checkout rate, sessions, avg tokens, avg duration, errors
- Funnel comparison matrix
- Performance table with totals
- By-store breakdown
- Eval configuration (stores, models, sequence turns)
- Recommendations
- Session replay links
The PDF is designed to be shared. Forward it to your dev team, include it in a client pitch, or attach it to a Jira ticket. Each session ID in the report links to a full replay.
Comparing Runs
Run the same collection weekly to track progress. Each run is versioned -- if you change the stores, models, or sequences, the version increments. Compare run-over-run to see if your store's checkout rate is improving or regressing.