Choosing Models
Test a store

Choosing Models

UCP Playground carries 26 models from 8 providers, each with different strengths. The right model depends on whether you are optimizing for completion rate, speed, or cost.

Available Models

The model picker is a grouped-by-provider menu, with a Compare Models button beneath it for side-by-side runs. The full lineup:

  • Anthropic -- Claude Opus 5, Claude Sonnet 5, Claude Opus 4.8, Claude Sonnet 4.6
  • OpenAI -- GPT-5.5, GPT-6 Luna, GPT-4o, o4-mini (Reasoning)
  • Google -- Gemini 3.1 Pro, Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 2.5 Pro, Gemini 2.5 Flash
  • xAI -- Grok 4.5, plus Grok 4.3 and Grok 3 Mini (Reasoning) (retired)
  • DeepSeek -- DeepSeek R1 (Reasoning), DeepSeek V3.2, DeepSeek V4 Pro, DeepSeek V4 Flash, DeepSeek V4.1 Flash
  • Moonshot AI -- Kimi K3
  • Alibaba -- Qwen 3.7 Max, Qwen 3.8 Flash, plus QwQ 32B (Reasoning) (retired)
  • Meta -- Muse Spark 1.3, Llama 4 Maverick, plus Llama 3.3 70B (retired)

Retired models (QwQ 32B, Grok 3 Mini, Llama 3.3 70B, Grok 4.3) keep their leaderboard stats and detail pages but can no longer start new runs. Prior-generation models (Claude Opus 4.8 / Sonnet 4.6, Gemini 3.5 Flash) stay selectable as lower-cost options.

Understanding the Trade-offs

Every model balances three factors: capability, speed, and cost (measured in token usage). More capable models tend to be slower and consume more tokens, but they handle complex shopping scenarios better -- especially variant selection, multi-step checkout flows, and recovery from unexpected tool responses.

Flagship Models

Claude Opus 5 is Anthropic's latest and most capable model. The Opus line excels at reasoning through complex catalogs, correctly matching variants (size, color, style), and navigating multi-step checkouts. The trade-off is higher token usage and longer session durations. Use it when you need the highest possible completion rate regardless of cost.

GPT-5.5 is a strong all-rounder that balances capability with efficiency. It handles most shopping scenarios well and uses fewer tokens than Opus. A solid default choice for production benchmarking.

Claude Sonnet 5, Grok 4.5, Gemini 3.1 Pro, Kimi K3, Qwen 3.7 Max, and Muse Spark 1.3 (Meta's multimodal reasoning model for long-running agentic work) are competitive flagship alternatives with strong tool-calling abilities. The most recently added models (the Claude 5 line, Kimi K3, Qwen 3.7 Max, Muse Spark 1.3, and the newest fast models below) are still accumulating Playground benchmark data — check the leaderboard for current numbers.

Balanced Models

Claude Sonnet 4.6 and Gemini 2.5 Pro sit in the sweet spot between capability and cost. They handle the full shopping funnel reliably for most stores while keeping token budgets reasonable. These are excellent starting points for initial testing.

Fast and Lightweight Models

The Gemini Flash line (Gemini 3.8 Flash, the newest, then 3.7, 3.6, 3.5, and 2.5 Flash) gives the fastest and cheapest options. They complete sessions quickly with minimal token usage, making them ideal for smoke tests, regression checks, and high-volume benchmarking. They may struggle with complex variant matching or stores that return unusual response formats.

GPT-6 Luna (OpenAI's low-cost GPT-6 tier), Qwen 3.8 Flash, DeepSeek V4.1 Flash, DeepSeek V4 Flash, DeepSeek V3.2, GPT-4o, and Llama 4 Maverick offer budget-friendly alternatives with varying strengths. They work well for straightforward catalogs and simple shopping queries.

Reasoning Models

DeepSeek R1 (Reasoning) exposes its chain-of-thought before responding. This makes it uniquely useful for debugging prompt injections and tool descriptions -- you can see exactly how the model interprets your instructions. Sessions use more tokens due to the reasoning overhead. Use it when you need to understand why the agent made a particular decision, not just what it did. o4-mini (Reasoning) is OpenAI's compact reasoning model, though it has struggled badly with MCP tool schemas in our benchmark -- check the leaderboard before relying on it.

Recommendations

If you are not sure where to start:

  • First-time testing -- start with Claude Sonnet 4.6 or GPT-5.5 for balanced performance and reliable results
  • Quick smoke tests -- use a Gemini Flash model (3.8, 3.6, or 2.5 Flash) for fast, inexpensive runs
  • Maximum completion rate -- use Claude Opus 5 when you need the best possible outcome
  • Comparing stores -- pick one model and use it consistently across stores so results are comparable
  • Debugging prompts -- use DeepSeek R1 (Reasoning) to see the model's reasoning and understand how it interprets tool descriptions
Tip

Check the Leaderboard to see real-world completion rates and average token usage for each model across hundreds of sessions. Data there reflects actual shopping performance, not synthetic benchmarks.