Will AI inference token costs keep falling, or is there a hard floor?
25 recorded positions from 17 people, first said Nov 15, 2023. They do not agree — the readings below are what each one actually argued.
Semiconductor improvement curve keeps driving cost down
Gustav Söderström · Dec 20, 2023 · hedged
AI cost reduction will keep following a Moore's-Law-like curve for a good while even if transistor density stalls, because neural hardware has barely started
Some say Moore's Law is ending for transistors per square inch, but the neural hardware curve is just beginning
Scope: for a good while; not necessarily via more transistors per square inch
15:42 20VC Roundtable: Spotify, Adobe & Linkedin CPOs on How AI Changes The Future of Product, Why AI is Now the Product, How TikTok Changed Product, Why Cost is the Biggest Barrier to LLM Usage & Why Incumbents Can Adopt AI Faster Than Any Prior Innovation Cyc
Andrew Feldman · May 26, 2026
Cost per unit of compute will keep falling substantially as every vendor's chips improve on tokens per unit time, per watt and per dollar
The history of the semiconductor industry is a massive reduction in cost per unit compute, and all players (Cerebras, Nvidia, AMD, Qualcomm, ARM) keep improving their designs
Scope: three to four year horizon
16:01 20VC: Cerebras CEO on the Future of Data Centres, Token Costs and Memory | We are Not in an Infra Bubble & Dario Got a Bad Deal with Elon for Compute | Should US Companies Sell to China & Why Most Layoffs are AI Washed with Andrew Feldman
Nikesh Arora · Jun 22, 2026
Older models were good enough in capability but inefficient on compute; efficiency gains are now arriving, so token prices will come down even though today's prices carry the R&D burden
The cost of R&D is currently being loaded onto what tokens have to pay for, while compute efficiency improves
26:10 20VC: Nikesh Arora on the Frontier Model Problem: Breadth vs Depth | The Future of Token Costs | Memory Becoming the Moat | Where Value Accrues: Infra, Models, or Apps? | Why Enterprise AI is Not Ready & Systems of Record vs Systems of Intelligence
Arvind Jain · Jul 11, 2026
Inference costs will fall by orders of magnitude because good technologists always figure out how to make technology very cheap.
That is the historical pattern for technology; the last six-to-nine months of per-token price increases is an anomaly nobody predicted.
Scope: framed as his bet/belief; acknowledges recent price rises are unexplained
32:38 20VC: Why OpenAI and Anthropic Won't Win the App Layer | Why Teams Will Get Bigger Not Smaller in a World of AI | Why AI Removes Incumbents Advantage of Bundling | China vs America: Who Wins the AI War with Arvind Jain, Co-Founder @ Glean
Measure cost per task not per token
Roman Chernin · Jun 8, 2026
The market is too obsessed with the nominal price of a GPU hour; total cost of ownership and platform-level optimization matter far more, and inference optimizations can change effective token cost by an order of magnitude
Effective uninterrupted runtime and tokens extracted per GPU vary hugely by platform quality, so the same $3-5 GPU price can produce completely different real costs for the customer
23:25 20VC: Nebius Co-Founder on AI Infrastructure Bubbles | The Real Impact of Open Source on OpenAI & Anthropic | How Price Elastic is Demand for Compute | Could Nebius Sell 10x More Compute If They Had It & more with Roman Chernin
Lin Qiao · Jul 20, 2026
Not all tokens are equal, so the industry should evaluate token economics per task rather than per token
Different models are more or less verbose; a model that is 2x cheaper per token but 2x more verbose costs the same to solve the same task
43:15 20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks
Harry Stebbings · Aug 10, 2026
A token is not a token — one provider can make a token go much further than another, and improving token efficiency is the provider's job
Lynn's analogy: you can drive around the whole block or drive straight to the store
9:13 20VC: Will OpenRouter Sell for $10BN to Stripe? | Why Chinese Open Models Are Beating America—and What Happens Next | Why Enterprises Are More Fearful of Anthropic and OpenAI Than China | Is the Routing Layer Becoming a Commodity with Alex Atallah
Jerry Murdock · Aug 22, 2026
A token is not a token — customization changes a token's value, and models differ in verbosity, which changes total cost
Companies like Fireworks help customize models; the more a model is customized the more its output and style change, and verbose versus terse models produce very different token volumes for the same job
Scope: may be roughly true today among frontier models
17:09 20VC: The AI Bubble Will Burst: Half the Neoclouds Will Die | China: Should We Ban Chip Exports & Be Fearful of Chinese Open-Source | Mag7: Who Dies and Who Thrives: Why Meta is Meh and Microsoft is Mega
Competition erodes the shortage premium until inference is a utility
Daniel Khachab · Oct 28, 2024
The foundation model layer is rapidly commoditizing with prices falling almost monthly, which is an advantage for the application layer
Choco's own inference price fell roughly 80% in six months without them doing anything, despite added transactions
16:29 20VC: Why SaaS is Dead | Why AI First Companies Will Win | We are in the Middle of a Cold War for AI Talent | Why Europe is F******* and We Need to Stop Whining with Daniel Khachab, Co-Founder @ Choco
Lin Qiao · Jul 20, 2026
Token costs will fall drastically, and cheaper infrastructure will cause usage to explode
Prices are high today only because of supply-chain shortage; in a free economy high prices invite entrants and competition, which drives cost down until inference becomes an affordable utility people stop thinking about
42:12 20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks
Labs hold token prices up to prove margins to capital markets
Nikesh Arora · Jun 22, 2026
Frontier model companies are currently value-maxing rather than token-maxing, and are keeping token prices high to show gross-margin profitability because capital markets won't keep funding another $100B of compute every year
Having raised at trillion-dollar valuations, they realize financial markets cannot bear repeated $100B compute rounds, so the only lever is charging more for their fastest-growing product — tokens
25:04 20VC: Nikesh Arora on the Frontier Model Problem: Breadth vs Depth | The Future of Token Costs | Memory Becoming the Moat | Where Value Accrues: Infra, Models, or Apps? | Why Enterprise AI is Not Ready & Systems of Record vs Systems of Intelligence
Harry Stebbings · Jul 11, 2026
Model providers raised per-token prices because they needed to demonstrate they were good businesses ahead of going public.
33:09 20VC: Why OpenAI and Anthropic Won't Win the App Layer | Why Teams Will Get Bigger Not Smaller in a World of AI | Why AI Removes Incumbents Advantage of Bundling | China vs America: Who Wins the AI War with Arvind Jain, Co-Founder @ Glean
Three levers hardware datacenter efficiency and algorithms drive costs down
Bret Taylor · Oct 2, 2024 · hedged
Inference costs are falling rapidly while quality rises, so margins for most AI use cases will improve and the cost of running AI could track something like Moore's law.
Tracked GPT model costs keep dropping as quality improves, and techniques like distillation plus hardware improvements let you serve smaller, cheaper, faster models at similar quality.
Scope: Moore's law is a trend, not a law; 'probably' on margin improvement
30:15 20VC: Bret Taylor: The AI Bubble and What Happens Now | How the Cost of Chips and Models Will Change in AI | Will Companies Build Their Own Software | Why Pre-Training is for Morons | Leaderships Lessons from Mark Zuckerberg
Andrew Feldman · Mar 24, 2025
The cost of inference will fall through three distinct levers: cheaper, higher-performance computers each generation, more efficient data centers with lower PUE, and much more efficient algorithms raising utilization
Cost is built from data center OpEx plus the cost of the computer plus algorithmic efficiency, and the industry gets better at all three over time, yielding more tokens per unit time for the same power
22:43 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman
Also on the record
Arthur Mensch · Apr 29, 2024 · hedged
Algorithmic efficiency, not falling hardware cost, is where the gains come from — roughly 100x algorithmic improvement in three years versus ~30% hardware cost reduction every two years
Compute cost reduces on the Nvidia roadmap but not faster than the workload grows, so Mistral bets on efficiency improvements
22:56 Algorithmic efficiency not hardware cost decline is the dominant driver roughly 100x vs 30 percent per two years
Clay Bavor · Jul 4, 2026
Token costs will have a floor because demand for frontier intelligence is effectively unbounded while GPU supply and energy are the rate limiter
Basic supply and demand: if the constraint is the number of Blackwells and H100s, and you must pay for energy and compute, price cannot fall below that input cost
13:53 Gpu and energy supply set a price floor
Clay Bavor · Jul 4, 2026
Open weights models are cheaper mainly because they avoid the margin stack of hosted frontier models, but the fundamental inputs — GPU capacity and power — remain constrained either way
The cost saving comes from cutting out hosting margin, not from removing the underlying compute and energy constraint
15:31 Open weights savings are margin only not compute
Clay Bavor · Jul 4, 2026
Running models locally or on device will not alleviate the server-side compute constraint, though it will make some consumer applications much better
Training needs petaflops to exaflops and inference needs a large burst of compute quickly; phones hit thermal limits, so frontier workloads can only run in data center racks of TPUs or GPUs
16:12 On device inference cannot relieve server compute
Anastasios Angelopoulos · Aug 3, 2026
Inference will get cheaper in the long run; specifically, once Anthropic goes public and its very high inference gross margins become public information, that transparency will exert downward pricing pressure
Markets become efficient over time, and full information about a vendor's margins gives buyers negotiating leverage they lack against a private company
27:11 Public margin disclosure gives buyers leverage to push prices down
Anastasios Angelopoulos · Aug 3, 2026 · hedged
A company with sufficiently dominant technology can sustain very high margins despite buyer knowledge, as Apple does
Dominant technology gives pricing power that information transparency doesn't erode
28:47 Dominant technology sustains high margins despite transparency
Des Traynor · Nov 15, 2023
Compute and LLM costs will fall and the market will stratify into tiers — fast versus accurate models, high-end versus simpler chipsets — as vendors chase every price point
No technology in history has failed to get cheaper; every piece of hardware or software is eventually undercut by a cheaper candidate serving more of the market
32:36 Falling costs create tiered markets of fast vs accurate models
Tom Hulme · Apr 10, 2025 · hedged
As models get lighter, more and more inference could happen on device at the edge
70:33 Lighter models enable more inference to shift to the edge device
Stan Boland · Apr 10, 2025
Even with edge inference, a large amount of silicon will be required, and the huge growth in inference investment will continue for a while
Chain-of-thought reasoning and test-time compute mean token generation keeps rising, and inference investment has grown roughly 57x in the last year
70:41 Chain of thought and test time compute keep driving massive inference silicon demand despite edge inference
Scott Belsky · Dec 20, 2023
AI margins will improve as models improve, because the push for more performant models and the push for more cost-efficient models are usually the same effort
Performance and cost efficiency are parallel efforts toward the same outcome
15:07 Performance gains and cost efficiency gains are the same underlying model improvement effort
Andrew Feldman · Mar 24, 2025
Today's mostly all-to-all connected models waste enormous compute on connections that produce nothing, and sparsity approaches like MoE are only the first of many efficiency gains to come
Human models are not all-to-all connected; in a neural network layer every element connecting to every other is not how learning actually happens — some connections are valuable and some worthless, and MoEs already avoid presenting all weights to each token
24:06 Sparsity and future connection pruning techniques will keep cutting compute waste
Your assistant can query this graph directly — 25 positions here, 19,646 across the corpus. Add 996.fm over MCP.