UCP Checker
Which AI Models Can Actually Complete a Checkout? The Benchmark Just Refreshed to the Frontier

Which AI Models Can Actually Complete a Checkout? The Benchmark Just Refreshed to the Frontier

The question the UCP Playground leaderboard exists to answer is a moving target: which AI models can actually complete an autonomous purchase — discover a product, build a cart, clear checkout, and place a real order at a real store, with no human after "go"? Every few weeks the frontier moves and that question needs re-asking. This week it moved a lot, so the benchmark refreshed to match: the lineup now spans 20 models across 8 providers, including the current flagship generation and two providers we'd never benchmarked before.

Here's what changed, and the two decisions behind it that make this more than a version bump.

The new arrivals

The current frontier is now on the board, each running as its own model so it earns its own record from its first session:

  • Claude Opus 5 and Claude Sonnet 5 (Anthropic) — the headline additions; a full major-version step on the models most people reach for first.
  • Grok 4.5 (xAI) — xAI's latest flagship, added as its own model (the retired Grok 4.3 keeps its history but is no longer selectable).
  • Gemini 3.6 Flash (Google) — the latest fast-tier Gemini.
  • DeepSeek V4 Flash — rounding out the open-weight fast tier alongside DeepSeek V4 Pro.

And two genuinely new providers join the benchmark for the first time:

  • Kimi K3 (Moonshot AI)
  • Qwen 3.7 Max (Alibaba)

That's the broadest agentic-checkout benchmark we've run — frontier, mid-tier, fast, and open-weight classes side by side, from Anthropic, OpenAI, Google, xAI, Moonshot AI, Alibaba, DeepSeek, and Meta.

Decision 1: we kept the previous generation on purpose

Refreshing a benchmark usually means retiring the old to make room for the new. We didn't. Opus 4.8, Sonnet 4.6, and the earlier Gemini Flash models stay live — as lower-cost options — because a shipped flagship doesn't make last generation worse, it makes it cheaper.

That matters because "which model completes a checkout most reliably" is only half the question a merchant or platform actually asks. The other half is "…at what cost per completed purchase?" Once agentic commerce runs at volume, every checkout carries a token bill, and the right model is rarely the most expensive one — it's the cheapest one that clears the task reliably. Keeping the prior generation on the board turns the leaderboard from a pure capability ranking into a price-performance one, which is the decision the market will actually be making.

Decision 2: every model runs under its own identity

A benchmark is only worth citing if its numbers mean what they say. Every model on the leaderboard runs under its own distinct identity — its own dataset, its own earned record. When a new version ships we add it as a new model and sunset the old one; we never quietly swap a newer model into an existing slot, because that would silently merge two different models' results under one name and corrupt the record. Retired models keep the history they earned; new models start clean.

That discipline is also why you'll see "metrics pending" on the new arrivals for a little while. We don't publish a score for a model we haven't measured. If you want to know how Opus 5 stacks up against Grok 4.5 or Kimi K3 at completing a real purchase, the honest answer this week is "give it a few days of sessions" — and then the leaderboard will tell you, measured rather than guessed. Watch the new models climb, or stumble, as the runs accumulate.

See it

The refreshed lineup is live on the UCP Playground leaderboard. Pick a model, pick a store, and watch it shop — and if there's a model you want benchmarked, tell us.

About UCP Checker

UCP Checker is the independent validation and observability layer for the Universal Commerce Protocol. We measure agentic commerce from two vantages: we crawl and grade every public UCP manifest (what a store declares), and through UCP Playground we run real agent-driven checkouts (what actually happens at the transaction layer). The model leaderboard is the runtime half — measured the same way for every model, without picking winners.

Check your domain's UCP status

See if your storefront is ready for agentic commerce in seconds.

Weekly UCP Report

Get the agentic commerce digest every Monday

Real adoption data, ecosystem trends, new spec versions, and the stores that broke or recovered this week. Read by founders and engineers building the next generation of commerce.

Free forever No spam Unsubscribe anytime

View a sample report →

Weekly UCP Report
Issue #50 · Sep 14, 2026
+664
new verified stores
Verified rate
Latest spec
Cart capability