Nvidia's inference share is reported as both rising and shrinking. The metric being skipped explains why.
Merchant GPU sales and hyperscalers' self-built chips are different markets. Most 2026 chip reports quietly conflate them into one inference number.
Two numbers, one chip market
Two reports came out within weeks of each other in 2026, both measuring Nvidia's inference market share, and they told opposite stories. One tracked Nvidia's share of inference silicon rising from 66% to 74% over a year, evidence, its authors argued, that the company's moat was widening rather than narrowing. Another, tracking custom ASIC shipments from Google, Amazon, Microsoft and Meta, showed those chips growing 44.6% year over year against 16.1% for merchant GPUs, nearly three times the rate, concentrated specifically in inference workloads.
Neither number is wrong. They are counting two different markets that both happen to use the word "inference," and most of the coverage repeating one figure or the other never says so.
A quick baseline: where 92% came from
The comparison most coverage reaches for is Nvidia's overall data-centre AI accelerator share, which analysts put at roughly 92% in 2023, during the period when nearly every AI workload, training and inference alike, ran on whatever GPU a team could get its hands on. That figure has since drifted down to somewhere around 80–85% across the whole data-centre accelerator market by revenue. It is a real decline, but a shallow one, and it blends training with inference, and merchant sales with a growing pool of chips that were never for sale.
The 66%-to-74% figure is a narrower, more recent cut of the same picture: inference specifically, tracked over the last year rather than since 2023. Reading the two side by side without noticing they cover different scopes and different time windows is how a single company's chip business ends up described, correctly, as both consolidating and fragmenting in the same set of headlines.
Where Nvidia's inference market share is genuinely rising
The 74% figure comes from tracking chips sold on the open market: GPUs that cloud providers, enterprises and smaller AI labs buy because they need inference or training capacity and don't have the scale, the in-house chip design team, or the multi-year lead time to build anything of their own. In that market, Nvidia's newest generation, Vera Rubin, keeps the company ahead on raw throughput and, more importantly, on the CUDA software stack that nearly every AI team already builds against.
Nobody buying compute on the open market today is choosing a slower, less-supported chip to save a percentage point of cost. That part of the story is real, and it is not shrinking.
The market that is actually moving
The other number measures something that barely existed at scale three years ago: chips hyperscalers design themselves and never sell to anyone. Google's TPU v7 ("Ironwood"), Amazon's Trainium 3, Microsoft's Maia 200 and Meta's MTIA are built to run one company's own inference traffic, at a volume high enough that the non-recurring engineering cost of a custom chip pays for itself.
None of those chips show up in a "units sold" column, because they are not sold. They are consumed internally, by the company that designed them, for workloads that company already controls end to end. A market-share tracker built around shipment revenue will correctly show Nvidia dominant, because it still sells more chips to more buyers than anyone else. It will just be blind to the largest and fastest-growing slice of compute that was never for sale in the first place.
Why the split shows up in inference and not training
Training a large model still rewards general-purpose flexibility: architectures change between projects, batch sizes vary, and a GPU fleet that can be repointed at the next experiment is worth its cost even when it isn't perfectly efficient for any single job. Inference is the opposite kind of workload once it runs at hyperscaler volume. The model is fixed, the request shape is known, and the same computation repeats billions of times a day.
That is exactly the profile a custom ASIC is built for: give up flexibility, get a large efficiency gain on the one job you actually run. Inference now accounts for roughly two-thirds of all AI compute, which is precisely why the custom-silicon share of it, not the training share, is where the two reports actually disagree.
The quiet winners aren't Nvidia's usual rivals
AMD is the company most coverage names as Nvidia's inference challenger, and it has picked up real share, an estimated 5 to 7%, driven by MI300X and MI325X adoption. The bigger shift in dollar terms is happening one layer down the supply chain, at the two companies that actually design the custom chips hyperscalers now use: Broadcom and Marvell.
Between them they handle the overwhelming majority of that design work, turning each hyperscaler's chip ambitions into physical silicon. A hyperscaler moving inference workloads off Nvidia GPUs and onto its own ASIC is not handing that business to AMD. It is handing the design contract to whichever partner built the chip, and keeping the compute itself in-house.
This is also why AMD's gains and the hyperscalers' gains don't cannibalise each other the way Nvidia's losses to either one might suggest. AMD's MI300X and MI325X compete for the same merchant-market buyers Nvidia sells to: cloud customers who want a second source and a lower price on a chip they can actually purchase. Custom ASICs compete for a different budget entirely, one that never went to the merchant market in the first place because the buyer was always going to build rather than buy once its own inference volume justified the cost. Both are real pressure on Nvidia. They are not the same pressure, and they don't add together into a single "share lost" figure the way a simple market-share chart implies.
| Dimension | Merchant GPU (Nvidia) | Hyperscaler custom ASIC |
|---|---|---|
| Who can buy it | Anyone with a cloud account or purchase order | Nobody — kept for the designer's own workloads |
| Where it wins | Training and variable workloads, or anyone below hyperscaler volume | Fixed, high-volume, repetitive inference at one company’s own scale |
| 2026 shipment growth | ~16.1% | ~44.6% |
| Software ecosystem | CUDA, the default across most AI tooling | Proprietary, built and maintained by the one company using it |
“Two reports about the same chip market can both be right if they are counting two different chips.”
What this changes if you are buying inference compute
For nearly every team outside the four or five companies large enough to design their own silicon, none of this changes the buying decision. You are still buying Nvidia GPUs, through a cloud provider or directly, because a custom-chip project needs a volume and an engineering budget that don't exist outside the hyperscalers.
What does change is the pricing environment underneath that decision. As Google, Amazon, Microsoft and Meta move a growing share of their own inference off merchant GPUs and onto chips they built themselves, they need relatively fewer Nvidia GPUs for internal use, even as they keep selling Nvidia-backed instances to customers. That is a second-order effect worth watching in cloud GPU pricing and availability over the next few quarters, separate from any question about whether Nvidia's headline market share is "safe."
It also changes what's worth asking a cloud provider before signing a multi-year inference commitment. A workload that is fixed, high-volume and predictable, the kind a hyperscaler would eventually justify building custom silicon for, is exactly the kind that provider may quietly be shifting onto its own chips behind an API that still says "GPU instance" on the invoice. Asking which silicon actually serves a given inference endpoint, and whether that choice can change without notice, is a newer question than it would have been three years ago, and one procurement teams are only beginning to add to vendor reviews.
None of this is a story about Nvidia losing. Nvidia's newest architectures keep shipping into a market that, on the open side, still has nowhere else credible to go. It is a story about a second market growing up next to the first one, large enough now that describing "the AI chip market" as a single number stopped being accurate somewhere in 2026, even though most coverage, and most procurement conversations, haven't caught up to that yet.
The two competing numbers about Nvidia's inference share will keep showing up in different reports for as long as one market sells chips and the other builds them for itself. The reconciliation isn't coming, because there is nothing left to reconcile.
Frequently asked questions
Related reading
AI Agent Liability Insurance Quietly Disappeared From Standard Policies in 2026
US insurers spent early 2026 rewriting general liability policies to exclude AI agent losses by default. Here is what changed, and what it takes to buy the cover back.
Cloud repatriation is 21% of workloads, not an exodus. 37signals and GEICO show what actually moves back.
Trade coverage claims 80% of companies are leaving the cloud. Flexera's actual 2026 data says 21%. 37signals and GEICO, at wildly different scales, show which workloads clear the bar.
Outcome-based AI pricing charges per resolution. Vendors decide what a resolution is.
Outcome-based AI pricing promises vendors only get paid for results. But the definition of 'resolved' is written by the vendor, and silence-as-consent clauses can bill an unresolved, frustrated customer as a win.