CI/CD Integration
Automate

CI/CD Integration

Run evals on every deploy to catch regressions before they reach production. Use the headless API with GitHub Actions, GitLab CI, or any CI system.

How It Works

  1. Create a collection with your stores, models, and sequences (one-time setup)
  2. Trigger a run via POST /api/v1/collections/{id}/run
  3. Poll for completion via GET /api/v1/collection-runs/{run_id}
  4. Check results -- assert on checkout rate, error count, or duration thresholds

GitHub Actions Example

name: UCP Store Evals
on:
  deployment_status:
    types: [completed]

jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - name: Run collection
        run: |
          RUN_ID=$(curl -s -X POST \
            -H "Authorization: Bearer ${{ secrets.UCP_TOKEN }}" \
            https://ucpplayground.com/api/v1/collections/${{ vars.COLLECTION_ID }}/run \
            | jq -r '.run_id')

          # Poll until complete (max 5 minutes)
          for i in $(seq 1 30); do
            STATUS=$(curl -s \
              -H "Authorization: Bearer ${{ secrets.UCP_TOKEN }}" \
              https://ucpplayground.com/api/v1/collection-runs/$RUN_ID \
              | jq -r '.status')
            [ "$STATUS" = "complete" ] && break
            sleep 10
          done

          # Check checkout rate
          RATE=$(curl -s \
            -H "Authorization: Bearer ${{ secrets.UCP_TOKEN }}" \
            https://ucpplayground.com/api/v1/collection-runs/$RUN_ID \
            | jq '.summary.checkout_rate')

          echo "Checkout rate: $RATE%"
          if (( $(echo "$RATE < 80" | bc -l) )); then
            echo "::error::Checkout rate $RATE% below 80% threshold"
            exit 1
          fi

API Endpoints

EndpointDescription
POST /api/v1/collectionsCreate a collection
GET /api/v1/collectionsList collections
POST /api/v1/collections/{id}/runTrigger a run
GET /api/v1/collection-runs/{id}Get run status + summary
GET /api/v1/collection-runs/{id}/pdfDownload PDF report
Tip

Store your collection ID as a CI variable and your API token as a secret. The collection persists across runs -- you only create it once and trigger runs on each deploy.

Scheduled Runs

For ongoing regression detection, set a cron schedule on the collection — runs dispatch automatically while is_active is true:

curl -X PATCH -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  https://ucpplayground.com/api/v1/collections/{id} \
  -d '{ "schedule": "0 9 * * 1" }'  # Every Monday at 9am

Invalid cron expressions are rejected at the API. Check past runs via GET /api/v1/collections/{id}/runs to track trends over time. Prefer triggering from your own CI instead? A scheduled GitHub Actions workflow calling POST /api/v1/collections/{id}/run works just as well.

Completion Webhooks

Set a webhook_url (HTTPS only) on the collection to get a POST when each run completes — no polling needed:

curl -X PATCH -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  https://ucpplayground.com/api/v1/collections/{id} \
  -d '{ "webhook_url": "https://ci.example.com/hooks/ucp" }'

The response includes a webhook_secret — shown once, rotated whenever the URL changes. Each delivery is signed with X-Playground-Signature: sha256=<HMAC-SHA256 of the raw body>; verify it before trusting the payload. The body carries event: collection_run.completed, the run id and status, and the full run summary (checkout rates, per-store and per-model breakdowns, errors). Failed deliveries retry twice with backoff.

Assertions

Set thresholds in the collection config and the server evaluates them when each run completes — the run summary carries passed_assertions (also in the webhook payload), so CI gates on one boolean:

"config": {
  "assertions": { "min_checkout_rate": 80, "max_avg_tokens": 50000, "max_errors": 0 }
}

Supported: min_checkout_rate, min_search_rate, max_avg_tokens, max_avg_duration_ms, max_errors. Unknown keys are rejected at the API. Per-assertion detail (threshold vs. actual) is in summary.assertions.

What to Assert On

  • checkout_rate >= 80 -- your baseline should be at least 80% for production-ready stores
  • errors.total == 0 -- zero tool errors means clean MCP implementation
  • avg_duration_ms < 30000 -- sessions over 30s may indicate MCP endpoint performance issues
  • avg_tokens < 50000 -- high token usage signals bloated product descriptions or inefficient agent behavior