The State of Agentic Commerce — June 2026
It's been five months since UCP was announced — from a launch-day reveal to a live commerce protocol with thousands of stores running it.
Last month the story was a slope, not a staircase: UCP grew without a platform flipping a switch, and we asked whether May's first-adopters — non-Shopify platforms, AP2, third-party namespaces — would survive contact with month two. June's answer: they survived, and most of them roughly doubled. The directory crossed 8,000 verified stores — a +56% month, the largest single jump since launch — but the more telling movement was at the edges: every deep primitive that sat at one or two stores in May (AP2 mandate, identity linking, buyer consent) is still tiny, and still doubled. For the first time since we started this report, the deep end of the spec showed a pulse — and the Technical Council is busy specifying the payment- and identity-layer extensions those primitives will need. The growth is real; the more interesting signal is that the protocol is starting to mature past "agent-shoppable."
This is the fifth monthly state-of-the-ecosystem report from UCP Checker. Here's what the data says as of June 13, 2026.
The numbers
- 8,259 verified UCP stores (up from 5,294 in May, +56%)
- 9,252 total domains tracked
- 3,251 new merchants discovered this month; 787 this week alone
- 8,230 verified stores on the latest v2026-04-08 spec (99.6%)
- ~94% of verified stores at A grade on UCP Score
Two straight months of ~30% growth became one month of +56%. Some of that is the discovery-rate step-up we shipped in May still compounding; most is real new supply — 3,251 genuinely new merchants entered the directory in 30 days, the most new merchants we've recorded. The shape from prior months holds: February was discovery, March expansion, April the Shopify spec migration, May the first slope. June is the slope steepening — and, for the first time, the deep end of the spec showing a pulse.
Shopify's head start, five months in
Five months in, the single most durable fact in this dataset is unchanged: Shopify is ~99% of the verified directory. The +56% came almost entirely from the Shopify long tail filling in, not from a new platform shipping a wave. Every non-Shopify platform combined — Custom & Headless (~58 verified), BigCommerce (4), WooCommerce (4), PrestaShop (2) — is still under 70 verified stores. SFCC, Adobe Commerce, Wix, and Squarespace have still not shipped a platform-level integration — the highest-impact event that keeps not happening. The one-platform structure is now five months old, and the question for July is the same as it was for May: who, if anyone, ships next.
The Custom & Headless cohort reads differently once you look at what's actually in it: a large share isn't stalled production stores but pre-production testing — developers pointing local environments through tunnels at our conformance and test harness on UCP Checker, and exercising real agent flows against their own builds on UCP Playground, validating against the spec before they ever ship a public manifest. That's the build loop working as intended, not a graveyard of failed attempts: when a platform hands you the boilerplate you compound; when you hand-build, you lean on the harness to reach a passing UCP Score before launch. The distance between attempts and verified is a tooling gap, not a spec problem — and that's exactly the gap the Score and the test harness exist to close.
UCP Score: the Lighthouse for UCP
Every web developer knows what a Lighthouse score is. You run it, you get a number out of 100 across a handful of categories, and you iterate against the breakdown until the number is green. It became the default because it was free, open, ungated, and ran the same way for everyone — so the score itself became the shared language for "is this page good enough."
UCP Score is the Lighthouse for UCP — the same idea for agentic commerce. A free, open, 0–100 readiness grade across three categories (Discovery, Conformance, Capability), run identically against every store, that turns the vague question "is my manifest agent-ready?" into a concrete, category-by-category checklist. We built it to be the neutral readiness meter for the protocol, and over five months it's become exactly that in practice: developers iterate against the Score breakdown the way they iterate against Lighthouse — we've watched a failing manifest climb to an A in the space of an afternoon, the developer re-running the Score between fixes. ~94% of the verified directory now carries an A; the score methodology is public, the per-domain score is ungated, and the whole verified dataset ships under CC-BY.
That's deliberate. A protocol needs one neutral yardstick everyone can point at — measured for everyone, owned by no one. UCP Score is that yardstick: the lighthouse you run to find out whether an agent can actually work with your store before you ship it.
Capability coverage: the cliff develops a pulse
| Capability | Verified adopters |
|---|---|
dev.ucp.shopping.checkout |
8,234 |
dev.ucp.shopping.order |
8,214 |
dev.ucp.shopping.fulfillment |
~8,215 |
dev.ucp.shopping.cart |
8,209 |
dev.ucp.shopping.catalog.* |
~8,200 |
dev.ucp.shopping.discount |
~8,185 |
| — the cliff — | |
dev.ucp.common.identity_linking |
11 |
dev.ucp.shopping.buyer_consent |
6 |
dev.ucp.shopping.ap2_mandate |
3 |
dev.ucp.shopping.checkout.embedded |
2 |
dev.ucp.shopping.payment |
0 |
The pattern is the same one we've reported since March — the core shopping capabilities ship as a Shopify-side bundle (~8,200 adopters each), then a ~750× cliff — but the cliff itself moved this month for the first time. Every deep primitive roughly doubled off its tiny base: identity linking 6 → 11, buyer consent 3 → 6, AP2 mandate 1 → 3 (catyai.io, houseofparfum.nl, ucp.travel). Payment capability: still 0.
Three stores is not adoption. But it's the first time the deep end has shown a consistent direction rather than a single fixed data point — and AP2 mandate, the primitive that makes an agentic transaction auditably user-authorised, is the rung that turns "agent-shoppable" into "agent-payable." When demand for AP2 turns into pressure — regulators, the working group's eventual requirements — this number moves the way checkout did once Shopify bundled it. The Technical Council is already laying that groundwork: a backlog of payment- and identity-layer extensions (split payments, 3DS2, delegated identity, per-channel buyer consent) is moving through review right now, even though live adoption is still three stores. The spec is being built ahead of the demand. Until it arrives, "UCP store" still means "agent-shoppable," not "mandate-credentialed."
Transports and payment handlers: the monoculture, slightly less alone
| Transport | Verified declarations |
|---|---|
| MCP | ~8,200 |
| Embedded | ~8,180 |
| REST | ~87 |
| A2A | 3 |
MCP and Embedded are universal because Shopify declares both. REST is the non-Shopify hand-build signal and it grew with the long tail — from 47 in May to ~87. A2A — Google's Agent2Agent transport, now governed by the Linux Foundation, formally added to UCP in v2026-04-08 — ticked from 2 to 3. Still the agent-native edge, not retail, but holding and growing.
Payment handlers tell the same monoculture story we flagged in February: Google Pay and Shop Pay are declared by ~100% of verified stores — the shared Shopify-managed handler IDs — and everything else is a rounding error. The payment partner roster on the registry (Stripe, Adyen, PayPal, Affirm, Splitit, Checkout.com and the card networks) is mature on paper; the live handler declarations are two Shopify-managed IDs and a handful of dev experiments. The gap between the partner roster and the live declarations is the clearest "declared vs. live" gap in the dataset.
How agents actually perform
The numbers above tell you which stores have UCP. This section is which stores work when an agent shops them. UCP Playground Evals crossed 2,000 recorded agent sessions this month — 2,089 end-to-end agent shopping runs across 232 unique stores and 16 frontier models, totalling ~110M tokens, run by many developers.
Outcomes: where the agent stops
| Outcome | Sessions | Share |
|---|---|---|
checkout_reached |
633 | 30.3% |
search_only (browsed, didn't cart) |
601 | 28.8% |
failed (provider error, refusal, max turns) |
535 | 25.6% |
cart_created (carted, didn't proceed) |
264 | 12.6% |
purchase_completed |
46 | 2.2% |
info_provided |
10 | 0.5% |
The dataset doubled and the shape held — with one sharpening. We now track all the way to purchase_completed, and the deepest funnel number is the one that matters: of the 633 sessions that reached checkout, only 46 completed a purchase — about 7%. Across all sessions, 2.2% end in a completed buy. Agents find products nearly everywhere (search works), build carts often, then stall: ~13% cart and don't proceed, ~26% fail outright. The categories of failure are the same ones we documented in UCP Variant Data: The #1 Reason Agent Checkouts Fail — variant mismatch, slow tokenisation, malformed cart responses, checkout redirect loops — and they're still where the distance between "has a manifest" and "an agent can buy from it" lives.
Model leaderboard
The field by recorded session volume this month — the live leaderboard breaks out checkout-conversion, cart, and speed per model:
| Model | Share of runs | Avg latency | Vendor |
|---|---|---|---|
| Gemini 3 Flash | 16.7% | ~19 s | |
| Claude Sonnet 4.5 | 16.5% | ~36 s | Anthropic |
| Claude Opus 4.6 | 12.5% | ~31 s | Anthropic |
| Gemini 2.5 Flash | 8.8% | ~13 s | |
| Gemini 3.1 Pro | 6.7% | ~46 s | |
| GPT-4o | 6.4% | ~20 s | OpenAI |
| GPT-5.2 | 6.2% | ~42 s | OpenAI |
| Llama 3.3 70B | 5.4% | ~40 s | Meta |
| Gemini 2.5 Pro | 5.3% | ~33 s | |
| Grok 4 | 3.9% | ~55 s | xAI |
| DeepSeek V3.2 | 3.9% | ~50 s | DeepSeek |
The field is Gemini- and Claude-heavy at the top — the four most-run models are two Google, two Anthropic — and the reasoning-tuned models (QwQ, Grok 3 Mini, o4-mini, DeepSeek R1) sit at the bottom of the run-count, consistent with the finding we've reported twice: shopping rewards decisive, sequential tool-calling, not deliberation. Per-model checkout-completion rates are on the live leaderboard.
The reliability gap, one more time
We've made this the editorial spine of every one of these reports, and June doesn't let us retire it — it sharpens it. By conformance the directory is in good shape (~94% A). But the funnel above is the counter-argument: a near-universally A-graded directory still converts an agent-reached checkout to a completed purchase only ~7% of the time. Conformance is not agent-readiness, and UCP Score doesn't claim to grade the second thing.
A clean schema doesn't tell you whether the cart endpoint accepts the variant the agent picked, whether response-time budgets hold under load, whether payment-handler tokenisation completes inside the agent's timeout, or whether the checkout URL drops the agent into an auth loop. UCP Playground is the harness developers use to exercise that second layer — and by design it surfaces failure modes, not steady-state performance. But the categories it surfaces are real, and they're what separate an A-graded manifest from a store an agent can transact against in production. The protocol's first phase was getting the schema right; the ecosystem did that. The next phase is the unglamorous second-order work — error recovery, schema robustness, response-time SLAs, variant-data hygiene — and the open question is whether that posture spreads from the engineering teams already running the loop to the long tail still on bundled defaults.
Spec and ecosystem
AP2 moved to the FIDO Alliance. Google's agent-payment mandate standard is now governed by FIDO — a sign the agent-payments layer is being organised at the standards level even though live adoption (three stores) is only just beginning.
The payment- and identity-layer pipeline is filling. The Technical Council spent late May moving exactly the extensions that would populate the capability cliff: a split-payments extension (#409) cleared for approval, 3D Secure (3DS2) support (#421) in deep-dive review, a returns extension (#257) under a phased plan, delegated identity providers (#423) flagged as critical for merchant flows, a loyalty extension (#340), and a buyer-consent V2 redesign moving from a single boolean to fine-grained reverse-DNS segments (consent.email, consent.sms) with submission restricted to the transaction-complete call (#407). The deep primitives are at three stores today because the spec for them is being written right now.
WebMCP is becoming a real transport. The TC scheduled a deep-dive to review Shopify's internal WebMCP prototype — the clearest signal yet that WebMCP, the browser-native transport, moves from "next" toward a measurable surface. We're tracking it; today it's still zero in the wild.
A v2026-04-08 patch is queued. The council aligned on bundling critical fixes — a schema correction and the identity-linking foundation — into a single backported patch and community announcement, rather than waiting for a full minor release. v2026-04-08 remains current; the next minor more likely lands late summer.
Tooling and governance. Shopify shipped a UCP CLI (mid-May) that wraps protocol semantics and transport so developers and agent builders can query the merchant ecosystem directly. And the council began planning vertical expansion — domain-specific tech councils and SDK/process changes as UCP moves beyond retail into travel, booking, and other verticals.
What to watch in July
Does the cliff keep its pulse? June was the first month the deep primitives (identity, AP2, consent) moved in the same direction. July's diagnostic: do they double again, or was June a one-off of vendor test deployments?
Does the payment layer get specced? The TC has split payments, 3DS2, delegated identity and buyer-consent V2 in flight. July's question is whether any graduate from review into the spec — giving the deep-end primitives somewhere to land.
WebMCP in the wild. With Shopify prototyping it and the TC reviewing it, the question is whether the first WebMCP-declaring store shows up in the crawl.
The platform wave that keeps not coming. SFCC, Adobe Commerce, Wix, Squarespace — any platform-level UCP integration is still the highest-impact possible event, and still hasn't happened in five months.
Sources
All data is from the UCP Checker crawler (re-checks every tracked domain at least every 24 hours) and UCP Playground's eval sessions, as of June 13, 2026. The verified-merchant dataset is published monthly on Hugging Face under CC-BY 4.0; the same data, a public REST API, the bulk checker, and the rest of our developer tools are ungated.
- Browse the directory: ucpchecker.com/directory
- Track adoption live: ucpchecker.com/stats
- Run a UCP Score: ucpchecker.com/score
- Model + store leaderboard: ucpplayground.com/evals
- UCP Technical Council minutes: meeting-minutes/tc/2026
- Previous report: State of Agentic Commerce — May 2026
Check your domain's UCP status
See if your storefront is ready for agentic commerce in seconds.
Get the agentic commerce digest every Monday
Real adoption data, ecosystem trends, new spec versions, and the stores that broke or recovered this week. Read by founders and engineers building the next generation of commerce.

