In-Chat Checkout Promised a Million Merchants. Eight Months Later, About 30 Were Live.
Walmart's own conversion data explains why AI shopping agents are winning at discovery and losing badly at the register.
What "In-Chat Checkout" Actually Delivered by February 2026
OpenAI launched Instant Checkout inside ChatGPT on September 29, 2025, with Etsy live on day one and a promise that Shopify's roughly one million merchants would follow. Walmart joined in November, listing about 200,000 products for in-chat checkout without leaving the chat window. The pitch was simple: skip the redirect, skip the cart page, let the agent finish the transaction where the conversation already was.
By February 2026, Forrester analyst Emily Pfeiffer counted how many of those million-plus eligible Shopify merchants had actually gone live on in-chat checkout. About 30. Not 30,000. Thirty.
That gap is the part most coverage of AI shopping agents skips on the way to a bigger claim about disruption. The infrastructure got built on schedule. The merchants mostly didn't show up for it, and it turned out, neither did the purchase behaviour.
Walmart's Own Numbers Explain the Gap
On March 4, 2026, Walmart EVP Daniel Danker told a Morgan Stanley TMT conference that in-chat checkout was "a very temporary moment in time." That is an unusual thing for the retailer running the largest live deployment of the feature to say about it in public. The data that surfaced over the following two weeks explains why he said it.
ChatGPT referrals were genuinely valuable traffic for Walmart: new-customer acquisition ran at roughly twice the rate Walmart sees from a typical search referral. But when the purchase itself happened inside the chat window, conversion came in about three times worse than sending that same shopper to walmart.com to finish the transaction.
“AI shopping works right up until the moment of payment. At that exact moment it currently works worse than not using AI at all.”
Checkout Was Never Really a UI Problem
It is tempting to read a 2x-worse acquisition-to-3x-worse-conversion split as a UX gap that better prompt design or a slicker in-chat cart will close. The friction says otherwise. CAPTCHAs, 3-D Secure step-up authentication, address verification, and fraud scoring were all built assuming a human is sitting at the keyboard, making the same judgment calls a fraud model expects a human to make. An agent hits every one of those checkpoints as a stranger.
The panel at Fortune's Brainstorm Tech conference on June 12, 2026 spent most of its time on the layer underneath the UX: who is liable when the agent gets it wrong. Courtney Robinson of Akoya put it plainly: liability is "wide open right now and being negotiated company to company," with no industry standard yet for who eats the cost of a fraudulent or mistaken agent-initiated purchase. Norman Menz of Flare warned that agents do not reduce ecommerce fraud risk, they multiply it, because a hijacked or spoofed agent can transact at machine speed with none of the friction that used to slow a human fraudster down. Matt Maher of M7 Innovations raised what he called perceptual liability: a shopper does not distinguish "my agent bought the wrong size" from "the merchant charged me wrong," and the refund complaint lands on the merchant either way, regardless of what the terms of service say.
None of that is a chatbot design problem. It is an unresolved question about who holds the risk, and until an answer exists, rational merchants have a reason to keep the last step, the one where money actually moves, on infrastructure they control.
The Protocols Shipped Anyway. They Pointed at Discovery Instead.
2026 was not short on payment infrastructure for agents. The Agentic Commerce Protocol (ACP), built by OpenAI with Stripe, handled the Instant Checkout transactions. Google shipped the Agent Payments Protocol (AP2). The Universal Commerce Protocol (UCP), launched by Google and Shopify on January 11, 2026 with more than 20 backers, added real-time catalog sync and multi-item carts in a March 19 update. A fourth standard, x402, tackled agent-to-agent payment at the protocol level.
What is notable is what none of that spring 2026 infrastructure push was actually optimizing for: completing more transactions inside a chat window. Shopify's clearest move came on March 24, when it switched on "Agentic Storefronts" by default for 5.6 million stores, making their catalogs cleanly readable by ChatGPT, Copilot, Google AI Mode, and Gemini. That is a discovery and listing mechanism, a way for an agent to find and describe a product accurately, not a checkout mechanism. The industry kept building rails for agents to find and describe things well. It mostly stopped building rails for agents to buy things directly.
The Merchants Actually Making Money From This
Two examples from Shopify's own merchant base show where the revenue is actually landing. Omnilux, which sells red-light therapy masks, attributed 3.2% of its March 2026 revenue to AI channels. Cozy Earth, a luxury bedding brand, reported 20x year-over-year revenue growth from AI channels. Neither number came from in-chat checkout. Both came from AI tools discovering the product through a structured, accurate feed and sending the shopper to the merchant's own site to complete the purchase. Shopify's vendor data puts structured feeds at converting roughly twice as well as feeds an AI tool has to scrape off a page.
The demand side backs this up. A Semrush survey of 1,030 US shoppers in December 2025 found 22% had ever completed a purchase directly inside an AI tool, 50% had bought something after using AI to research it first, and 69% expected AI to play a larger role in their shopping going forward. That is not a checkout replacement. It is a research and discovery habit that occasionally, not routinely, closes the loop in-chat, which lines up with what Walmart and Shopify were both seeing on the merchant side by March.
| Dimension | In-chat checkout (ACP-style) | Discover-in-AI, buy-on-site |
|---|---|---|
| Where the transaction completes | Inside the chat window | Merchant's own site |
| Fraud and refund liability | Unresolved, negotiated per deal | Sits with merchant's existing stack |
| Walmart-measured conversion | ~3x worse than redirect | Baseline |
| Engineering investment needed | New checkout/payment integration | Structured, real-time product feed |
| 2026 evidence of revenue | About 30 live Shopify merchants | Omnilux 3.2% of revenue; Cozy Earth 20x YoY |
What to Build This Quarter, If You Run a Storefront
None of this means agentic commerce stalled. Shopify reports AI-driven traffic up roughly 7x and AI-attributed orders up roughly 11x since January 2025; those figures are vendor-reported and unaudited, but the direction matches every other data point here. What stalled specifically is the last six inches of the transaction, the part where money changes hands inside someone else's chat window instead of on the merchant's own domain.
For a team deciding where to spend engineering time this quarter, the data points at one order of operations. Build a structured, real-time product feed compatible with UCP and Shopify's Agentic Storefronts before building anything that looks like conversational checkout: that is where Omnilux and Cozy Earth's numbers actually came from. Keep transaction completion on infrastructure you control, at least until a liability standard exists for agent-initiated purchases; that is not AI caution, it is the exact choice Walmart made in March after running the largest live test anyone has. Track AI referral as an acquisition metric, separate from a conversion metric, so a channel that is quietly doubling your new-customer rate does not get judged, or funded, by a checkout number it was never going to win.
Every panelist at the Fortune event in June agreed on the one open question that actually matters here: nobody has settled who pays when an agent gets it wrong. Until that is resolved, discover in AI, buy on your own site is not a compromise merchants are settling for. It is the equilibrium the data has already found.
Frequently asked questions
Related reading
AI Agent Benchmarks Got Gamed to Near-Perfect Scores Without Solving a Single Task
Eight major AI agent benchmarks hit 73-100% scores without an agent solving the underlying task. A second 2026 study found the same gap honestly: a 37% lab-to-production drop and a 50x cost swing.
AI Agents Are Now Provisioning 80% of New Databases. The Review Process Didn't Scale With Them.
AI agents now provision most new databases on platforms like Neon. The 80% figure is a velocity number, not a governance one, and the failures showing up are schema drift and orphaned branches, not bad SQL.
AI coding agents didn't fix the bottleneck in software delivery. They moved it to code review.
AI coding agents increased developer throughput in 2026, and median code review time along with it, up 441% in one large dataset. Four independent studies show where the extra work actually goes.