Will the AI inference provider layer consolidate under hyperscalers or remain a competitive marketplace?
15 recorded positions from 9 people, first said Feb 17, 2025. They do not agree — the readings below are what each one actually argued.
Undifferentiated neoclouds die in a shakeout
Harry Stebbings · Jun 15, 2026
Nebius's co-founder's core realization is that although huge amounts of money want raw capacity and compute, a long-term sustainable business requires building a full-stack product
37:10 20VC: Micron Will Be More Valuable Than Meta | How Export Controls Helped Not Hurt China | Power is the Bottleneck to AI | Why Dario Has Done a Disservice to AI with his Labour Replacement Messaging with Aravind Srinivas, Founder @ Perplexity
Jerry Murdock · Aug 22, 2026 · hedged
At least half of today's neoclouds will disappear within thirty-six months
Many are undifferentiated and dependent on cheap capital; a dislocation would wipe out the weakest immediately
Scope: within 36 months; accelerated by an economic disruption
0:00 20VC: The AI Bubble Will Burst: Half the Neoclouds Will Die | China: Should We Ban Chip Exports & Be Fearful of Chinese Open-Source | Mag7: Who Dies and Who Thrives: Why Meta is Meh and Microsoft is Mega
Compute supply access decides which inference providers survive
Steeve Morin · Feb 24, 2025 · hedged
If forced to buy one stock today it would still be NVIDIA, because of supply
NVIDIA has the supply; AMD and others may become the buy once switching software ships
Scope: 'today at least'; expects to change his answer to AMD/Tenstorrent if his own software ships
44:27 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML
Anjney Midha · Apr 14, 2026
Which four or five inference companies win will be determined by access to compute supply
Compute is the input to the product — without it there is nothing to sell, like building a steam engine without coal
46:29 20VC: Anj Midha on Investing $300M into Anthropic | The Early Days of Anthropic & How 21 of 22 VCs Turned it Down | The Four Bottlenecks to Compute | What the China Has Smashed and Why We Should Be Worried
Nvidia prevents hyperscaler gpu monopoly
Andrew Feldman · May 26, 2026 · hedged
Nvidia has deliberately funded, backstopped and over-allocated to neoclouds in order to create competitors to the traditional hyperscalers, producing a dependence that is probably unhealthy
Reads Nvidia's allocation behaviour toward neoclouds as a strategy to counterbalance the hyperscalers
Scope: framed as 'I think'; dependence characterized as 'probably not healthy'
0:00 20VC: Cerebras CEO on the Future of Data Centres, Token Costs and Memory | We are Not in an Infra Bubble & Dario Got a Bad Deal with Elon for Compute | Should US Companies Sell to China & Why Most Layoffs are AI Washed with Andrew Feldman
Alex Atallah · Aug 10, 2026
Hyperscalers will not be able to buy up all the GPUs and drive independent inference providers out of business, because NVIDIA deliberately avoids customer concentration and wants a heterogeneous compute market
One of NVIDIA's top priorities is avoiding customer concentration; they want many customers with separate allocations and competition at the compute layer, which is good for the ecosystem, for NVIDIA and for end users
7:42 20VC: Will OpenRouter Sell for $10BN to Stripe? | Why Chinese Open Models Are Beating America—and What Happens Next | Why Enterprises Are More Fearful of Anthropic and OpenAI Than China | Is the Routing Layer Becoming a Commodity with Alex Atallah
Also on the record
Aravind Srinivas · Jun 15, 2026
Neocloud/inference-and-datacenter companies can plausibly reach $10B revenue and $100B valuations, but only if open source models stay within roughly twelve months of the frontier
If the open-source-to-frontier gap widens to 15-18 months, these companies have no business model beyond renting capacity to OpenAI or Anthropic
37:42 Neocloud viability hinges on open source staying near frontier
Steeve Morin · Feb 24, 2025
Buying NVIDIA GPUs inside a hyperscaler's cloud is a losing game because stacked margins (TSMC ~60%, NVIDIA ~90%, cloud ~30%) leave the customer a thin crust on a big cake; you should keep optionality across chip providers
each layer of the supply chain takes a large margin, so going all in on one provider leaves almost no economics for you
16:00 Stacked margins across the compute supply chain favor keeping optionality across chip vendors
Jonathan Ross · Sep 29, 2025
Buying capacity from GPU cloud providers like CoreWeave doesn't solve the compute problem because every such provider has a finite GPU allocation
Everyone renting out GPUs is working from a fixed allocation of chips
39:07 Renting neocloud capacity doesnt solve scarcity finite allocation everywhere
Jonathan Ross · Sep 29, 2025
Extreme customer concentration in AI compute means buyers will increasingly decide on what makes their own business successful rather than on brand, so alternative chips will get used
Roughly 35-36 customers account for 90-99% of total market spend, which gives them enough power to make their own decisions instead of defaulting to the safe brand
64:14 Concentrated buyer power forces brand agnostic chip adoption
Alex Atallah · Aug 10, 2026
The inference provider layer turned out to be a competitive marketplace rather than a hyperscaler monopoly, because independent providers host open weight models faster and better than the hyperscalers
Independent providers were much faster to host new models and to solve the edge cases of serving them; you never hear of people running GLM on a hyperscaler, they use Fireworks, Together and similar
6:01 Independent providers out execute hyperscalers
Matan Grinberg · Jun 13, 2026
Which neocloud provider wins between Nebius and CoreWeave doesn't matter; the desirable end state is one where application users don't know which compute provider is under the hood
Speaking from the application layer, where the underlying compute should be invisible
68:11 Compute provider should become an invisible commodity to apps
Andrew Feldman · May 26, 2026
Hyperscalers' bundled security, software layers and credibility are genuinely valuable to most enterprise segments, but for the segment that just wants cheap compute those same strengths become a cost disadvantage
Value has to be manufactured and comes at a cost; buyers who don't want the extras (leather seats vs Naugahyde) won't pay for them, so the market is segmented
13:07 Market segments by whether buyers want the hyperscaler bundle
Andrew Feldman · May 26, 2026
Companies that deploy their own hardware in their own data centers (Google, Cerebras) have a significant cost advantage over neoclouds, which must buy hardware carrying Nvidia's 80% gross margins and then layer their own margin on top
Neoclouds pay Nvidia's 80% gross margin and then need their own margin, stacking cost into the data center
17:54 Stacked vendor margins disadvantage neoclouds versus self deployers
Jonathan Ross · Feb 17, 2025
A competitor pricing below Groq is no longer a real threat
So much capital is flooding into AI that customers will run on Groq in order to lose less money, so undercutting doesn't win the business
39:56 Capital abundance blunts price undercutting threats for efficient providers
Your assistant can query this graph directly — 15 positions here, 19,646 across the corpus. Add 996.fm over MCP.