Is NVIDIA's GPU architecture, born from repurposed graphics hardware and scaled vertically, hitting fundamental physical and architectural limits for AI workloads?
9 recorded positions from 3 people, first said Nov 22, 2023. They do not agree — the readings below are what each one actually argued.
Gpus as repurposed graphics hardware break down at llm scale due to memory bottlenecks
Jeff Seibert · Nov 22, 2023
The top models are memory-bandwidth bound as well as compute bound, so faster GPUs alone won't unlock performance and chip companies face new challenges
OpenAI shared a tech talk describing how they tune their data centers around this constraint
34:25 20VC: Why OpenAI Will Become an Infrastructure Play, Why Apple Will Win in an AI World, Why Google is the Most Vulnerable Incumbent, Will LLMs Be Commoditised, Which Startups Are Thin vs Thick Wrappers on Top of LLMs with Jeff Seibert, Founder @ Digits
Steeve Morin · Feb 24, 2025
GPUs are a clever repurposing of graphics hardware for parallel matrix work rather than a dedicated AI architecture, and that approach starts to break down for LLMs
GPUs were designed to render pixels in parallel; twenty years ago people tricked them into doing parallel compute (GPGPU), but LLMs are so big that the constant memory transfers become the bottleneck
Scope: breaks down specifically at LLM scale
13:27 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML
Andrew Feldman · Mar 24, 2025
The fundamental GPU architecture, with its off-chip memory, is poorly suited to inference and can be beaten there — and NVIDIA knows it
Off-chip memory is the wrong architectural choice for inference workloads
Scope: GPUs will nonetheless continue to do well in inference
0:00 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman
Andrew Feldman · Mar 24, 2025
In AI workloads the arithmetic itself is trivial; the genuinely hard problem is moving results and intermediate results to and from memory and between GPUs
The core operation is matrix multiplication and an FMAC any second-year EE student can build, but there are a huge number of them and the data has to be constantly moved and split across memory and chips
6:09 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman
Andrew Feldman · Mar 24, 2025
The most underrated threat to NVIDIA's dominance is that the GPU's fundamental architecture with off-chip memory is not well suited to inference
Off-chip memory is a structural weakness for inference workloads
Scope: NVIDIA will still continue to do well in inference
53:09 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman
Also on the record
Steeve Morin · Feb 24, 2025
NVIDIA is pursuing a very vertical, foot-to-the-floor GPU scaling strategy that is running into physical limits
With Blackwell they assembled two chips whose combined surface was so big it started to bend, breaking heat sink contact, and they've pushed to a thousand watts requiring liquid cooling
12:28 Vertical gpu scaling hits physical limits like chip bending and thermal constraints
Andrew Feldman · Mar 24, 2025
The rise of AI would place a different kind of pressure on processors — specifically on memory bandwidth and communication structure — creating an opening to build a better-suited machine
A genuinely new workload means the demands AI software makes of the underlying processor differ from prior workloads, which is a computer architect's opportunity
4:47 Novel ai workload demands open an opening for new processor architectures
Andrew Feldman · Mar 24, 2025
The market still uses HBM because, short of going to wafer scale, there was no credible alternative — GPUs were always built that way for graphics
Off-chip memory was part of the GPU's original advantage over CPUs; now that dedicated AI chips exist, that former advantage has become their weakness
10:13 Hbm persists only for lack of a credible wafer scale alternative
Andrew Feldman · Mar 24, 2025
NVIDIA is aware of the memory limitation but is constrained by it because it buys rather than makes memory, and only three to five suppliers exist
NVIDIA depends on SK Hynix, Samsung and Micron for memory, so its options are limited and the choice is part of a complex architectural trade-off that has otherwise served it extremely well
10:44 Nvidia cannot fix its memory limits because it depends on third party memory suppliers
Your assistant can query this graph directly — 9 positions here, 19,646 across the corpus. Add 996.fm over MCP.