Skip to content

Debates

Is NVIDIA's GPU architecture, born from repurposed graphics hardware and scaled vertically, hitting fundamental physical and architectural limits for AI workloads?

9 recorded positions from 3 people, first said Nov 22, 2023. They do not agree — the readings below are what each one actually argued.

Gpus as repurposed graphics hardware break down at llm scale due to memory bottlenecks

Jeff Seibert · Nov 22, 2023

The top models are memory-bandwidth bound as well as compute bound, so faster GPUs alone won't unlock performance and chip companies face new challenges

OpenAI shared a tech talk describing how they tune their data centers around this constraint

34:25 20VC: Why OpenAI Will Become an Infrastructure Play, Why Apple Will Win in an AI World, Why Google is the Most Vulnerable Incumbent, Will LLMs Be Commoditised, Which Startups Are Thin vs Thick Wrappers on Top of LLMs with Jeff Seibert, Founder @ Digits

Steeve Morin · Feb 24, 2025

GPUs are a clever repurposing of graphics hardware for parallel matrix work rather than a dedicated AI architecture, and that approach starts to break down for LLMs

GPUs were designed to render pixels in parallel; twenty years ago people tricked them into doing parallel compute (GPGPU), but LLMs are so big that the constant memory transfers become the bottleneck

Scope: breaks down specifically at LLM scale

13:27 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML

Andrew Feldman · Mar 24, 2025

The fundamental GPU architecture, with its off-chip memory, is poorly suited to inference and can be beaten there — and NVIDIA knows it

Off-chip memory is the wrong architectural choice for inference workloads

Scope: GPUs will nonetheless continue to do well in inference

0:00 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Andrew Feldman · Mar 24, 2025

In AI workloads the arithmetic itself is trivial; the genuinely hard problem is moving results and intermediate results to and from memory and between GPUs

The core operation is matrix multiplication and an FMAC any second-year EE student can build, but there are a huge number of them and the data has to be constantly moved and split across memory and chips

6:09 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Andrew Feldman · Mar 24, 2025

The most underrated threat to NVIDIA's dominance is that the GPU's fundamental architecture with off-chip memory is not well suited to inference

Off-chip memory is a structural weakness for inference workloads

Scope: NVIDIA will still continue to do well in inference

53:09 20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Also on the record

Steeve Morin · Feb 24, 2025

NVIDIA is pursuing a very vertical, foot-to-the-floor GPU scaling strategy that is running into physical limits

With Blackwell they assembled two chips whose combined surface was so big it started to bend, breaking heat sink contact, and they've pushed to a thousand watts requiring liquid cooling

12:28 Vertical gpu scaling hits physical limits like chip bending and thermal constraints

Andrew Feldman · Mar 24, 2025

The rise of AI would place a different kind of pressure on processors — specifically on memory bandwidth and communication structure — creating an opening to build a better-suited machine

A genuinely new workload means the demands AI software makes of the underlying processor differ from prior workloads, which is a computer architect's opportunity

4:47 Novel ai workload demands open an opening for new processor architectures

Andrew Feldman · Mar 24, 2025

The market still uses HBM because, short of going to wafer scale, there was no credible alternative — GPUs were always built that way for graphics

Off-chip memory was part of the GPU's original advantage over CPUs; now that dedicated AI chips exist, that former advantage has become their weakness

10:13 Hbm persists only for lack of a credible wafer scale alternative

Andrew Feldman · Mar 24, 2025

NVIDIA is aware of the memory limitation but is constrained by it because it buys rather than makes memory, and only three to five suppliers exist

NVIDIA depends on SK Hynix, Samsung and Micron for memory, so its options are limited and the choice is part of a complex architectural trade-off that has otherwise served it extremely well

10:44 Nvidia cannot fix its memory limits because it depends on third party memory suppliers

Your assistant can query this graph directly — 9 positions here, 19,646 across the corpus. Add 996.fm over MCP.