Should application-layer AI startups build in-house research capability or rely on foundation model providers?
39 recorded positions from 27 people, first said Apr 28, 2023. They do not agree — the readings below are what each one actually argued.
Own your intelligence rather than rent it
Clem Delangue · May 12, 2023
Using one model behind an API is faster and easier at first but carries more long-run risk, while training and optimizing your own models is what builds real capability and differentiation.
With an API you never build internal AI capability, cannot optimize the models so they stay inherently more expensive, and you end up undifferentiated versus competitors; it's like using Squarespace or Wix in the early web instead of writing your own code.
Scope: short-term ease versus long-term risk trade-off; applies to companies that want to truly differentiate
14:35 20VC: Why The Future of AI Is Open Not Closed, Why We Are Years Away From AI Being Autonomous, Why AI Founders Do Not Need to Move to the Valley & Why Founders Should Not Meet Investors in Between Rounds with Clem Delangue @ Hugging Face
Clem Delangue · May 12, 2023
Incumbents will win if AI just means calling AI APIs, but startups that actually train models, create new architectures and optimize models themselves have an opportunity to do things 10 to 100 times better than incumbents.
Building AI as a radical paradigm switch — rather than consuming APIs — is really hard for incumbents to do.
Scope: conditional on what one means by 'AI startup'
23:37 20VC: Why The Future of AI Is Open Not Closed, Why We Are Years Away From AI Being Autonomous, Why AI Founders Do Not Need to Move to the Valley & Why Founders Should Not Meet Investors in Between Rounds with Clem Delangue @ Hugging Face
Clem Delangue · May 12, 2023
All companies will eventually have their own AI models — their own ChatGPT or GPT-4
41:40 20VC: Why The Future of AI Is Open Not Closed, Why We Are Years Away From AI Being Autonomous, Why AI Founders Do Not Need to Move to the Valley & Why Founders Should Not Meet Investors in Between Rounds with Clem Delangue @ Hugging Face
Anton Osika · Aug 18, 2025
The step-function change for Lovable would be AI that has much more context about who it is talking to and hyper-personalizes its guidance — and Lovable will have to solve that itself, both via its agentic chain and eventually by training models
Model providers won't deliver this specific personalization; it requires Lovable's own agentic architecture and, over time, world-class model training talent
Scope: would require spending on the order of $100M to hire people who train models
20:12 20VC: Lovable CEO Anton Osika on $120M in ARR in 7 Months | The Honest Truth About Defensibility and Unit Economics for AI Startups | The State of Foundation Models: Long Grok, Short OpenAI, Why | Replit vs Lovable vs Bolt: What Happens
Nick Frosst · Sep 1, 2025
The model layer and the application layer are not really separable — to build the best product on an LLM you should be training the model for that product
LLMs generalize well but not as well as people assume, so the best model for a given interface is one trained on that interface
14:23 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst
Lin Qiao · Jul 20, 2026
Companies with meaningful AI traffic in production are concluding that owning their intelligence beats renting it
Once you scale, optimization kicks in; optimization requires control, and control requires building on top of open models and turning your own data into intelligence
Scope: applies to companies past the pilot stage and into scaling
56:41 20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks
Lin Qiao · Jul 20, 2026
Within the next three years every company will own its own intelligence, and this will be a must-have rather than optional
By analogy to software: every company builds its own software stack because each solves a unique problem and wants full control, choosing which layers to build and which are commodity knowledge not worth building
Scope: companies will still pick and choose which parts of the stack to build vs. use off the shelf
73:21 20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks
Harry Stebbings · Aug 10, 2026
Companies should own their intelligence rather than rent it, and will have specialized models trained on their own proprietary data
11:01 20VC: Will OpenRouter Sell for $10BN to Stripe? | Why Chinese Open Models Are Beating America—and What Happens Next | Why Enterprises Are More Fearful of Anthropic and OpenAI Than China | Is the Routing Layer Becoming a Commodity with Alex Atallah
Alex Atallah · Aug 10, 2026 · hedged
Agent labs have a strong incentive to build their own models and distribute them through their agents, and this trend is only just beginning (Cursor already has one, Lovable does not publicly)
Companies known for making agents have a clear incentive to make and distribute their own models; Cursor already has one, Lovable doesn't yet, and Jeff Dean is starting an agent lab out of Google
25:50 20VC: Will OpenRouter Sell for $10BN to Stripe? | Why Chinese Open Models Are Beating America—and What Happens Next | Why Enterprises Are More Fearful of Anthropic and OpenAI Than China | Is the Routing Layer Becoming a Commodity with Alex Atallah
Build deep in the stack except capital intensive pretraining
Gustav Söderström · Dec 20, 2023
There is no point in Spotify trying to compete with OpenAI on frontier models; Spotify should build only what the labs won't build and buy the rest
OpenAI's stated goal is AGI, which drives different incentives — it won't optimize for delivering two minutes of audio to half a billion users a day, for cost efficiency, or for pragmatic products
24:17 20VC Roundtable: Spotify, Adobe & Linkedin CPOs on How AI Changes The Future of Product, Why AI is Now the Product, How TikTok Changed Product, Why Cost is the Biggest Barrier to LLM Usage & Why Incumbents Can Adopt AI Faster Than Any Prior Innovation Cyc
Scott Belsky · Dec 20, 2023
A company should only build its own models where it is the best company in the world to build them, and partner for everything else
Adobe has the customer use-case understanding and the data such that no one should build a better imaging model, whereas building a massive LLM is better left to companies whose business that is
25:05 20VC Roundtable: Spotify, Adobe & Linkedin CPOs on How AI Changes The Future of Product, Why AI is Now the Product, How TikTok Changed Product, Why Cost is the Biggest Barrier to LLM Usage & Why Incumbents Can Adopt AI Faster Than Any Prior Innovation Cyc
Bret Taylor · Oct 2, 2024
Spending capital on pre-training is nonsensical for any company that isn't an AGI research lab or one of the behemoths — it's just burning capital
It's like an entrepreneur building their own data center by hand: it isn't what the company does, and you don't want a huge upfront investment while still searching for product-market fit, especially when high-quality fine-tunable models like GPT-4o mini and Llama 3.1 exist
Scope: unless you are an AGI research lab or one of the behemoths
19:15 20VC: Bret Taylor: The AI Bubble and What Happens Now | How the Cost of Chips and Models Will Change in AI | Will Companies Build Their Own Software | Why Pre-Training is for Morons | Leaderships Lessons from Mark Zuckerberg
Clay Bavor · Jul 4, 2026
Companies should be willing to invest as far down the technology stack as necessary to build the product they want
Google's early success came from building its own cluster architectures and distributed systems on commodity hardware; similarly, agent frameworks didn't exist and had to be invented from scratch
Scope: stops short of doing their own pre-training, leaving that capex to the labs
7:49 20VC: Open Models vs Frontier Models: Who Actually Wins? | The $100,000 Token Budget Every Engineer Will Need | Why Forward-Deployed Engineers Are the Future of Enterprise AI with Clay Bavor, Co-Founder of Sierra
Clay Bavor · Jul 4, 2026
The right strategy for application companies is to slipstream behind the labs' and hyperscalers' capital-intensive investments, taking as much off the shelf as possible while still engineering deeply where it matters
Capital-intensive areas are best left to those already spending there; Sierra instead fine-tunes proprietary models on top of open-weights models
Scope: applies to deeply capital-intensive areas
9:25 20VC: Open Models vs Frontier Models: Who Actually Wins? | The $100,000 Token Budget Every Engineer Will Need | Why Forward-Deployed Engineers Are the Future of Enterprise AI with Clay Bavor, Co-Founder of Sierra
Lin Qiao · Jul 20, 2026
Building on open models rather than training proprietary ones was the right bet, because open weights give the user full control to modify and build on the model
Once a model is released openly you own the weights and can change them however you want; openness gives control to the user, a principle rooted in their PyTorch experience
Scope: at founding it was a big bet since open models were in their infancy
13:17 20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks
Julien Bek · Aug 24, 2026
Cursor is the best agent company outside Sequoia's portfolio
They are in one of the most important markets, and while others dismissed AI wrappers, Cursor was first to understand you could post-train models and go deeper into the stack
68:28 20VC: Inside Sequoia's Investment Committee: Lessons from Don Valentine, Doug Leone and Alfred Lin | How the SpaceX and Citadel Deals Went Down | What Sequoia Specifically Looks for in Founders with Julien Bek
Rely on model providers hire applied engineers
Sarah Guo · Apr 28, 2023
There are fewer than ten instances where it makes sense for a company to train a very large model from scratch; the vast majority of companies should apply, fine-tune, or build on models others have made available via APIs or open source.
Training from scratch requires 20-30 researchers who know how to do it, 10,000+ GPUs and many months, which is extremely expensive, and most companies don't need it.
Scope: framed as her personal point of view; other parts of the stack are also viable
12:32 20VC: In AI Who Wins? Startups or Incumbents? What Happens to Wealth Inequality? Why Will $10BN+ Companies Only Have 10 People | Why Defensibility in Startups is BS & Speed is Everything? Why Large Groups Worsen Decision-Making with Sarah Guo
Emad Mostaque · May 17, 2023
Mid-size companies are not shut out of AI: they should not train their own models but can use open-source models or hire providers, and building on this technology is not hideously complicated
Everyone training their own model is like everyone launching their own university; implementing GPT-4 is not hard — the missing piece is design patterns and the intention to build and integrate, not capability
59:19 20VC: Why the AI Bubble Will Be Bigger Than The Dot Com Bubble, Why AI Will Have a Bigger Impact Than COVID, Why No Models Used Today Will Be Used in a Year, Why All Models are Biased and How AI Kills Traditional Media with Emad Mostaque, Founder & CEO @
Aravind Srinivas · Jun 5, 2024
Not training its own foundation models is a decisive advantage for Perplexity — training would have exhausted its funding and left nothing to acquire users
Competing in foundation model training requires committing to roughly three years of spend in advance to secure thousands of GPUs, whereas paying for APIs and inference leaves capital free to win users and lets the company benefit from any model commoditization
41:40 20VC: Perplexity's Aravind Srinivas on Will Foundation Models Commoditise, Diminishing Returns in Model Performance, OpenAI vs Anthropic: Who Wins & Why the Next Breakthrough in Model Performance will be in Reasoning
Ishan Mukherjee · Apr 4, 2025
AI application companies should not hire researchers — they should work with the best intelligence providers and hire pragmatic applied AI engineers who ship products
They're building applications and products, not doing research; centralizing focus on the model providers gets you to a great application faster
Scope: for application-layer companies
58:52 20Sales: How the Best Sales Teams Use AI to Win Enterprise Deals | Sales Teams Will Be Dramatically Smaller | How to Ramps Sales Reps Way Faster | Why Unpaid Design Partners are BS | Why this Generation of Sales is Soft with Ishan Mukherjee @ Rox
Mati Staniszewski · Sep 8, 2025
It is worth building and shipping products on top of outside research even when you have no internal research initiative behind them
Scope: a reversal of their prior rule of only doing product innovation backed by their own research
66:17 20VC: ElevenLabs Hits $200M ARR: The Untold Story of Europe's Fastest Growing AI Startup | The Real Cost of AI from Talent to Data Centres | How US VCs are in a Different League to Europeans | The Future of Foundation Models with Mati Staniszewski
Mike Cannon-Brookes · Oct 13, 2025
There will be multiple competing foundational models rather than one winner, so Atlassian should not train its own and should instead build world-class expertise at rapidly evaluating and shipping new models to customers
With a new model roughly every three months and four to six competitors, the durable capability is picking one up, testing it, deciding where it's better and deploying it fast; Atlassian is unlikely to be able to compete at training
Scope: stated as a bet that is 'proving' right, with a caveat that nobody can claim to be proven right in AI yet
19:57 20VC: Atlassian CEO on Why Everything is Overvalued & Are We in an AI Bubble | Do Margins Matter & Does Defensibility Exist in an AI World | Is Per Seat Pricing Dead & The Future of Vibe Coding with Mike Cannon-Brookes
Task shape decides in house versus frontier
Saam Motamedi · Jul 15, 2024 · hedged
Personal agents, horizontal enterprise agents and code generation are the use cases that require tight model-application integration, whereas customer service AI does not
In customer service the startup should focus on everything except the model
Scope: companies are taking both approaches in code generation and it's unresolved which wins
17:05 20VC: Why We Are in a Bubble & Now is Frothier Than 2021 | Why $1M ARR is a BS Milestone for Series A | Why Seed Pricing is Rational & Large Seed Rounds Have Less Risk | Why Many AI Apps Have BS Revenue & Are Not Sustainable with Saam Motamedi @ Greylock
Victor Riparbelli · Jan 15, 2025
A company should only build its own models where it can be the best in the world, which in practice means a narrow domain directly tied to a larger customer workflow with compounding advantages
You work backwards from what customers are trying to do; building only makes sense if you can be best-in-world and the model compounds
Scope: Synthesia's chosen narrow domain is humans presenting to camera and dialogue-driven content
34:03 20VC: Why Scaling Laws Will Not Continue | OpenAI vs Anthropic vs X.ai: Who Wins and Why | How Far Will Model Providers Go Into the Application Layer | The End State for Models: Many Specialised or Few Generalised with Victor Riparbelli @ Synthesia
Miles Clements · Mar 9, 2026
Cursor was right to build its own models, provided they are understood as specialized coding models rather than generalist models.
Specialized models serving professional coding tasks for enterprise users differentiate the product; they don't need to write poetry or explain apple pie recipes
Scope: scoped to specialist coding models, not generalist frontier models
10:08 20VC: Inside Accel's $4BN Growth Investing Machine | Cursor is Dead is Total BS: Here is Why | What Missing Rippling and ElevenLabs Taught Us | Are $2BN-$10BN IPOs Dead | Why Now is a Great Time to be Thoma Bravo with Miles Clements
Adam Foroughi · Apr 27, 2026
Recommendation and ad-ranking systems require custom-built models; you cannot defer to a general large language model to decide the next ad to show and get comparable performance.
This branch of machine learning hit its stride a decade ago and powers Facebook's, TikTok's and AppLovin's ad systems; a purpose-built model outperforms an LLM given the same user data.
33:50 20VC: Applovin: $160BN Market Cap, $5.48BN Revenue, $10M EBITDA Per Head | Why the Best Do Not Need Mentorship | Why Founders Should Not Angel Invest | Why Kindness in Business Will Slow You Down with Adam Foroughi
Shiv Rao · May 16, 2026
The build-versus-frontier decision should follow the shape of the problem: binary 'ring the bell' tasks are best served by in-house models, while problems you'll never be perfect at but where the market rewards continuous improvement should ride the frontier wave
For binary tasks an in-house model is faster, cheaper and can be set-and-forgotten so you can move on; for open-ended quality problems you want to keep inheriting frontier improvements
Scope: offered as a principle since 'nobody knows the answer'; healthcare has many use cases of the second kind
18:35 20VC: Lessons from Jensen Huang on "Founder Mode" | How to Know if OpenAI or Anthropic Will Kill your Company | How USV Liking Music Made Them $1BN on an Investment | The Five Year Desert to Product Market Fit & a $5.3BN Valuation with Shiv Rao @ Abridge
Vertical specialized models are built by application developers not offered off the shelf by model providers
Arthur Mensch · Apr 29, 2024
Vertical, task-specific models will be built by application makers rather than offered off the shelf by model providers
The only way to make a low-latency model that is very good at a specific task is to strip out the general-purpose aspect, since general models are bloated
14:14 20VC: Mistral's Arthur Mensch: Are Foundation Models Commoditising | How Do We Solve the Problem of Compute | Is There Value in the Application Layer | Open vs Closed: Who Wins and Mistral's Position
Arthur Mensch · Apr 29, 2024
Mistral's role is to provide tools that let developers create customized specialized models without expert AI knowledge
Making a specialized model is a very hard job tightly tied to how the pre-trained model was created, and AI expertise is scarce
14:53 20VC: Mistral's Arthur Mensch: Are Foundation Models Commoditising | How Do We Solve the Problem of Compute | Is There Value in the Application Layer | Open vs Closed: Who Wins and Mistral's Position
Also on the record
Anton Osika · Aug 18, 2025
It is too early for Lovable to optimize or adapt models to its specific application; better to iterate fast on what the AI can do
The AI is doing completely different things every month, so optimizing models for today's behavior would lock in work that becomes obsolete
12:52 Too early to adapt models to application iterate on capability instead
Shiv Rao · May 16, 2026
Abridge had to build its own models because frontier models couldn't deliver the latency required for in-workflow clinical use
Milliseconds matter for doctors who have no patience for new technology; high-stakes in-room workflows like approving orders or visit diagnoses before the patient leaves require insane performance and latency that frontier models couldn't achieve elegantly
20:39 Latency in workflow forces in house models
Amjad Masad · Apr 25, 2026
The opportunity to train your own model has reopened, because open source models are good and coding models are plateauing, letting companies fine-tune on their own data for their use case
Plateauing frontier performance plus strong open source means domain-specific fine-tunes can lead for a window
10:27 Plateau plus open source reopens the fine tuning window
Aaron Levie · May 22, 2024
Adopting the best external model is the only way an incumbent survives, no matter how good it believes its custom-trained model is or how much it fears handing data to a future competitor
You can wish your proprietary model were better, but if the external model is better for the product you're delivering to customers, refusing it is burying your head in the sand — the classic innovator's dilemma
28:57 Adopting the best external model is required for survival refusing is innovators dilemma
Emad Mostaque · May 17, 2023
Open owned models address a completely different TAM from proprietary API models, and companies will need both
Regulated companies can only send a limited amount of their data to OpenAI, so a model they own and train in their own cloud or on-prem plus GPT-4 access is the best of both worlds
14:26 Open owned models and proprietary apis serve different tams so companies need both
Harry Stebbings · Jul 27, 2025
It is a fundamental change to product that a company's output quality is now determined by third-party model releases outside its control
Lovable's quality jumped significantly when Claude 4 was released — the equivalent would be being told Monzo's product will get much better in Q3 or Q4 for reasons nothing to do with you
42:00 Product quality now hinges on external model releases outside company control
Reid Hoffman · Jun 10, 2024
Every cloud provider and every scale software company will need serious in-house AI capability and talent, though not necessarily frontier models
The highest-quality cognitive services will come from a blend of models, so everyone has to participate — via acquisition, refining open source, or building their own
10:41 Every company needs in house ai talent though not necessarily frontier models
Lin Qiao · Jul 20, 2026
Vertical application companies should own their own intelligence, because the bespoke harness that orchestrates tool calls needs to be co-trained with the model powering it
They hold proprietary workflow knowledge and data, and accuracy of deciding which tools to call is bespoke and must be trained jointly with the model
23:49 Bespoke tool calling harness must be co trained so own the model
Winston Weinberg · Jan 19, 2026
AI talent will matter again for application-layer companies, because differentiation will come from building custom solutions on newly available proprietary enterprise data rather than from the model alone
Application-layer companies started as wrappers where the model was the whole product; over time much of their differentiation comes from core software plus, now, proprietary-data problems that are genuine AI problems
53:38 Proprietary enterprise data work requires in house ai talent
Mati Staniszewski · Sep 8, 2025
In 2022 there was no real build-vs-leverage decision for model-dependent startups — existing models were plainly inadequate, so the only investor question was whether the team could build something better; that calculus is different today
Pre-ChatGPT there was almost no public attention on models and what existed in the market simply wasn't good
10:49 In 2022 inadequate models forced building today buying is viable
Your assistant can query this graph directly — 39 positions here, 19,646 across the corpus. Add 996.fm over MCP.