What drives the trend toward smaller AI models: task specialization or raw efficiency (speed and cost)?
9 recorded positions from 7 people, first said May 17, 2023. They do not agree — the readings below are what each one actually argued.
Efficiency not specialization drives smaller models
Christian Kleinerman · Sep 22, 2023
Both paths will advance — compute gets faster and cheaper while research pushes model compression — but smaller models are better
If you want one or more model calls in the serving path of a consumer product, you only have a budget of a few milliseconds, so size will matter
Scope: cites a startup working entirely on model compression for latency and cost
21:40 20VC: Are Foundation Models Becoming Commoditised? Do OpenAI and Anthropic of the World Have a Sustaining Moat? Why Smaller Models May Work Better? Why Incumbents with Data Power Win the AI War with Christian Kleinerman, SVP Product @ Snowflake
Arthur Mensch · Apr 29, 2024
Mistral 7B succeeded because it filled a missing spot in the performance-to-efficiency space, showing there was a lot of slack in compressing models
7B is the size that runs efficiently on a MacBook or smartphone, and prior 7B models weren't smart enough for interesting applications, so casual developers on laptops and gaming GPUs adopted it immediately
7:59 20VC: Mistral's Arthur Mensch: Are Foundation Models Commoditising | How Do We Solve the Problem of Compute | Is There Value in the Application Layer | Open vs Closed: Who Wins and Mistral's Position
Arthur Mensch · Apr 29, 2024
There is large market demand for efficiency rather than raw scale, so the right strategy is to hit top performance for a given cost and model size while also scaling up
The 7B reception showed strong interest in efficiency, which motivated Mixtral 8x7B and 8x22B
9:03 20VC: Mistral's Arthur Mensch: Are Foundation Models Commoditising | How Do We Solve the Problem of Compute | Is There Value in the Application Layer | Open vs Closed: Who Wins and Mistral's Position
Steeve Morin · Feb 24, 2025
RAG does not drive model sizes down; efficiency does — if you can hit the same performance with a smaller model, you will use the smaller one
Less is better: speed and efficiency are what pushes toward smaller models, not specialization
59:04 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML
Specialized smaller models outperform giant do anything models
Cristóbal Valenzuela · Aug 28, 2023
Larger parameter counts do make models better across more tasks and modalities, but the right model size depends on the task, and there is real opportunity in smaller models specialized to specific purposes
We're already seeing smaller task-specific approaches in the language domain
Scope: depends on what you're trying to do
16:21 20VC: Why AI Models are not a Moat, Where Does the Value in AI Accrue; Startups or Incumbents, What the World Has Got Wrong About AI, Why AI Needs a New Story and Who is the Right People to Tell it with Cris Valenzuela, Co-Founder & CEO @ Runway
Christian Kleinerman · Sep 22, 2023
Model size matters for generic consumer products but matters much less for specialized enterprise use cases, where a smaller fine-tuned model can outperform a generic large one
A consumer product like ChatGPT must know every topic and language so cumulative knowledge is valuable; in the enterprise, size mainly drives cost and latency, and there are plenty of examples of smaller fine-tuned models beating generic models on a specific purpose or dataset
Scope: use-case dependent
20:39 20VC: Are Foundation Models Becoming Commoditised? Do OpenAI and Anthropic of the World Have a Sustaining Moat? Why Smaller Models May Work Better? Why Incumbents with Data Power Win the AI War with Christian Kleinerman, SVP Product @ Snowflake
Bastian Lehmann · Apr 8, 2024
Figure is one of the standout AI companies, and smaller models trained for specific purposes are more promising than giant do-anything models
Figure's latest robots are remarkable, and he sees a lot of remarkable companies working with smaller, more dedicated models rather than 'a giant call center that can do anything'
51:10 20VC: Postmates Founder Basti Lehmann on How the Uber Deal Went Down and How a $2.65BN Deal Turned into $5BN, Why Great VCs Add No Value and VC Value Add is BS Marketing & Why The Biggest Companies in History Will be Born Today and Replace Incumbents
Model compression to preserve performance at smaller size is a key r and d frontier
Emad Mostaque · May 17, 2023
We are already past what was considered impossible: a single file of a few hundred gigabytes can pass every exam except English literature
Two years ago everyone said this was impossible, and there is no known floor to how small the models can get
Scope: nobody knows how low parameter counts can go
35:29 20VC: Why the AI Bubble Will Be Bigger Than The Dot Com Bubble, Why AI Will Have a Bigger Impact Than COVID, Why No Models Used Today Will Be Used in a Year, Why All Models are Biased and How AI Kills Traditional Media with Emad Mostaque, Founder & CEO @
Jeff Seibert · Nov 22, 2023
Compressing models to preserve performance at smaller size will be one of the most interesting areas of AI R&D
There is a counter-push against pure scale, and Apple Silicon shows you can push efficiency rather than raw compute when that is what the use case requires
Scope: for certain use cases
24:45 20VC: Why OpenAI Will Become an Infrastructure Play, Why Apple Will Win in an AI World, Why Google is the Most Vulnerable Incumbent, Will LLMs Be Commoditised, Which Startups Are Thin vs Thick Wrappers on Top of LLMs with Jeff Seibert, Founder @ Digits
Your assistant can query this graph directly — 9 positions here, 19,646 across the corpus. Add 996.fm over MCP.