Skip to content

Debates

Will AI token consumption keep compounding as inference costs fall?

15 recorded positions from 10 people, first said Mar 24, 2025. They do not agree — the readings below are what each one actually argued.

Usage compounds faster than inference costs fall

Andrew Feldman · Mar 24, 2025

The size of the inference market is the number of people using AI times how often they use it times how much compute each use takes, and all three factors are growing simultaneously right now

Training makes AI and inference consumes it; the rare simultaneous growth of users, frequency, and compute per use is what produces the current off-the-charts growth

17:21 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Andrew Feldman · Oct 6, 2025

Inference demand growth is the product of three simultaneously fast-growing variables — number of users, frequency of use, and compute per use — which is why its scale is so hard for people to intuit.

Human minds struggle with geometric growth, and here three multiplied variables are all compounding at once.

22:52 20VC: Cerebras CEO on Why Raise $1BN and Delay the IPO | NVIDIA Showing Signs They Are Worried About Growth | Concentration of Value in Mag7: Will the AI Train Come to a Halt | Can the US Supply the Energy for AI with Andrew Feldman

Shiv Rao · May 16, 2026 · hedged

Goldman's projection that agents will drive a 24x increase in token consumption over five years is directionally right, though the specific multiple is unknowable.

Usage compounds as you adopt the technology — even as inference costs fall you consume more, and for jobs you can never be good enough at, demand is a bottomless pit; companies are reorganizing so each person covers far more surface area.

Scope: no confidence in the exact 24x figure; directional only

41:09 20VC: Lessons from Jensen Huang on "Founder Mode" | How to Know if OpenAI or Anthropic Will Kill your Company | How USV Liking Music Made Them $1BN on an Investment | The Five Year Desert to Product Market Fit & a $5.3BN Valuation with Shiv Rao @ Abridge

Brendan Foody · Jun 1, 2026

Token consumption in the enterprise will keep rising very significantly before any leveling off, because cheaper and better models drive more total consumption rather than less

Jevons paradox — as models improve ~10x year over year and cost per performance falls, total model consumption goes up, the same dynamic as making humans more efficient leading to more jobs

42:42 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility, The Model is the Product | Token Spend Will Exceed Headcount Spend in 5 Years | The True Cost of Hiring AI Researchers in the Valley Today with Brendan Foody

Roman Chernin · Jun 8, 2026

Cheaper intelligence increases rather than reduces compute consumption — each efficiency gain lets people solve more complex tasks or make previously uneconomic tasks viable

The DeepSeek week when Nebius stock fell ~40% was simultaneously the best commercial sales week in company history, because customers realized inference economics finally worked in production

10:14 20VC: Nebius Co-Founder on AI Infrastructure Bubbles | The Real Impact of Open Source on OpenAI & Anthropic | How Price Elastic is Demand for Compute | Could Nebius Sell 10x More Compute If They Had It & more with Roman Chernin

Lin Qiao · Jul 20, 2026 · speculative

Token volumes could grow 20x to 100x by the end of next year

We are at a very early stage of the S-curve of explosion

Scope: 'could be possible'; wide range

37:17 20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks

Alex Atallah · Aug 10, 2026

Falling token prices are good for OpenRouter's business because usage grows more than proportionally, close to a perfect Jevons paradox

When OpenAI's Luna price dropped 10x on OpenRouter over two weeks, usage grew 13x and then stabilized, with few confounding variables despite DeepSeek and GLM launching at good prices

Scope: acknowledges no one has done a great job modeling Jevons paradox; this is a spot story

19:31 20VC: Will OpenRouter Sell for $10BN to Stripe? | Why Chinese Open Models Are Beating America—and What Happens Next | Why Enterprises Are More Fearful of Anthropic and OpenAI Than China | Is the Routing Layer Becoming a Commodity with Alex Atallah

Usefulness threshold crossed so daily use keeps demand climbing

Andrew Feldman · Mar 24, 2025

AI was a novelty rather than genuinely useful until late 2024, when it crossed into being part of everyday workflow for ordinary people outside Silicon Valley

Earlier models were 'cool' but people didn't know what to do with them; now a marketing team not using an LLM several times a day isn't doing its job, and his father, brothers and doctors use it

Scope: turning point placed in Q4 2024 running into 2025

18:12 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Andrew Feldman · May 26, 2026

Compute demand will not peak so long as AI keeps improving in usefulness, because usage is now daily, spreading to harder problems and across all demographic groups

Around 2025 the models got smart enough to be genuinely useful rather than a novelty, which flipped AI from something people tried to something people use every day — from 85-year-olds to 11-year-olds, not just 28-year-olds in Silicon Valley

Scope: conditional on frontier models continuing to get smarter and more useful

11:34 20VC: Cerebras CEO on the Future of Data Centres, Token Costs and Memory | We are Not in an Infra Bubble & Dario Got a Bad Deal with Elon for Compute | Should US Companies Sell to China & Why Most Layoffs are AI Washed with Andrew Feldman

Cheaper faster technology diffuses into new applications expanding demand

Andrew Feldman · Mar 24, 2025

Improvements in compute, algorithms and data will make AI faster and cheaper, and cheaper-faster technology causes entirely new applications to emerge

The history of computing: as computers got faster and cheaper they diffused into cars, pockets, dishwashers, TVs and kids' toys; diffusion of innovation accelerates when things get faster and cheaper

28:24 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Andrew Feldman · Mar 24, 2025

There are no examples in compute in fifty years where making things cheaper and faster shrank the market; the market always gets bigger

Scope: specific to compute/the tech industry

29:17 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Also on the record

Aravind Srinivas · Jun 15, 2026 · hedged

Token costs for a given task will fall — within roughly twelve months an open source model as good as today's frontier could be ~10x cheaper — so spend on today's workloads won't stay high

Open source catches up and inference gets cheaper; paired with the right harness and connectors, existing developer workflows will run fine on far cheaper models

19:34 Cheap open models collapse spend on todays workloads

Matan Grinberg · Jun 13, 2026 · hedged

There may be a short-term contraction in usage of the very frontier models as the ROI hangover hits, and that would be healthy

Better to correct with awareness than to stay blind to the spend and suffer a sudden shock later

19:57 Frontier usage contracts short term as roi scrutiny bites

Ryan Petersen · Jun 20, 2026

LLM spend at Flexport will rise sharply but eventually plateau: frontier models stay worth paying for in coding and product, while workflow automation shifts to good-enough open source, and once processes are automated there's no reason to keep spending.

Diminishing returns to frontier models on plain workflow automation, plus likely price deflation; the spend is a means to automating manual work, not an end.

18:45 Spend plateaus once the manual work is automated

Harry Stebbings · Feb 21, 2026

We will end up with inference running twenty four hours a day across every single thing we do

8:52 Always on continuous inference across everything we do

Your assistant can query this graph directly — 15 positions here, 19,646 across the corpus. Add 996.fm over MCP.