CI/CD Integration
Run evals on every deploy to catch regressions before they reach production. Use the headless API with GitHub Actions, GitLab CI, or any CI system.
How It Works
- Create a collection with your stores, models, and sequences (one-time setup)
- Trigger a run via
POST /api/v1/collections/{id}/run - Poll for completion via
GET /api/v1/collection-runs/{run_id} - Check results -- assert on checkout rate, error count, or duration thresholds
GitHub Actions Example
name: UCP Store Evals
on:
deployment_status:
types: [completed]
jobs:
eval:
runs-on: ubuntu-latest
steps:
- name: Run collection
run: |
RUN_ID=$(curl -s -X POST \
-H "Authorization: Bearer ${{ secrets.UCP_TOKEN }}" \
https://ucpplayground.com/api/v1/collections/${{ vars.COLLECTION_ID }}/run \
| jq -r '.run_id')
# Poll until complete (max 5 minutes)
for i in $(seq 1 30); do
STATUS=$(curl -s \
-H "Authorization: Bearer ${{ secrets.UCP_TOKEN }}" \
https://ucpplayground.com/api/v1/collection-runs/$RUN_ID \
| jq -r '.status')
[ "$STATUS" = "complete" ] && break
sleep 10
done
# Check checkout rate
RATE=$(curl -s \
-H "Authorization: Bearer ${{ secrets.UCP_TOKEN }}" \
https://ucpplayground.com/api/v1/collection-runs/$RUN_ID \
| jq '.summary.checkout_rate')
echo "Checkout rate: $RATE%"
if (( $(echo "$RATE < 80" | bc -l) )); then
echo "::error::Checkout rate $RATE% below 80% threshold"
exit 1
fiAPI Endpoints
| Endpoint | Description |
|---|---|
POST /api/v1/collections | Create a collection |
GET /api/v1/collections | List collections |
POST /api/v1/collections/{id}/run | Trigger a run |
GET /api/v1/collection-runs/{id} | Get run status + summary |
GET /api/v1/collection-runs/{id}/pdf | Download PDF report |
Store your collection ID as a CI variable and your API token as a secret. The collection persists across runs -- you only create it once and trigger runs on each deploy.
Scheduled Runs
For ongoing regression detection, set a cron schedule on the collection — runs dispatch automatically while is_active is true:
curl -X PATCH -H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
https://ucpplayground.com/api/v1/collections/{id} \
-d '{ "schedule": "0 9 * * 1" }' # Every Monday at 9amInvalid cron expressions are rejected at the API. Check past runs via GET /api/v1/collections/{id}/runs to track trends over time. Prefer triggering from your own CI instead? A scheduled GitHub Actions workflow calling POST /api/v1/collections/{id}/run works just as well.
Completion Webhooks
Set a webhook_url (HTTPS only) on the collection to get a POST when each run completes — no polling needed:
curl -X PATCH -H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
https://ucpplayground.com/api/v1/collections/{id} \
-d '{ "webhook_url": "https://ci.example.com/hooks/ucp" }'The response includes a webhook_secret — shown once, rotated whenever the URL changes. Each delivery is signed with X-Playground-Signature: sha256=<HMAC-SHA256 of the raw body>; verify it before trusting the payload. The body carries event: collection_run.completed, the run id and status, and the full run summary (checkout rates, per-store and per-model breakdowns, errors). Failed deliveries retry twice with backoff.
Assertions
Set thresholds in the collection config and the server evaluates them when each run completes — the run summary carries passed_assertions (also in the webhook payload), so CI gates on one boolean:
"config": {
"assertions": { "min_checkout_rate": 80, "max_avg_tokens": 50000, "max_errors": 0 }
}Supported: min_checkout_rate, min_search_rate, max_avg_tokens, max_avg_duration_ms, max_errors. Unknown keys are rejected at the API. Per-assertion detail (threshold vs. actual) is in summary.assertions.
What to Assert On
- checkout_rate >= 80 -- your baseline should be at least 80% for production-ready stores
- errors.total == 0 -- zero tool errors means clean MCP implementation
- avg_duration_ms < 30000 -- sessions over 30s may indicate MCP endpoint performance issues
- avg_tokens < 50000 -- high token usage signals bloated product descriptions or inefficient agent behavior