Is retrieval-augmented generation (RAG) a durable solution for injecting external data into models, or fundamentally limited by unsolved chunking and context-window constraints?
10 recorded positions from 7 people, first said Apr 21, 2023. They do not agree — the readings below are what each one actually argued.
Rag is a step change solution for hallucination and model customization
Douwe Kiela · Jun 30, 2023
Retrieval augmented generation is the right architecture for enterprise language models because decoupling memory from generative capacity solves hallucination, attribution, updateability and privacy at once
Grounding generations in retrieved memory reduces hallucination, gives attribution for free, lets you add/remove/revise information on the fly, compresses compute into the memory for efficiency, and cleanly separates data plane from model plane for privacy guarantees
Scope: framed for enterprise first-principles design
7:00 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI
Richard Socher · Aug 18, 2023
Retrieval augmentation for large language models is far more important than most people appreciate
41:06 20VC: Does Value Accrue to Incumbents or Startups in the AI Race, Why Model Size Matters More Than Data Size, Why Artificial General Intelligence is Far Away, Why Carpenters Will Be Paid More Than Software Engineers & Future of Jobs with Richard Socher
Aidan Gomez · Aug 19, 2024
Retrieval augmented generation is a step change for hallucination and a game changer for customising models
Querying a knowledge base lets the model cite sources so answers can be audited, reduces the need to make things up because it has reference material, and lets the model answer questions about private data like your email inbox that isn't in the public web
37:24 20VC: Chips, Models or Applications; Where is the Value in AI | Is Compute the Answer to All Model Performance Questions | Why Open AI Shelved AGI & Is There Any Value in Models with OpenAI Price Dumping with Aidan, Gomez, Co-Founder @ Cohere
Decoupling data plane from model plane preserves privacy while retaining control
Tomasz Tunguz · Apr 21, 2023 · hedged
The dominant AI architecture will be one where the model is deployed to the customer's own data account, executed next to the data, and then pulled back out with the result
This mirrors the existing bifurcation where the data plane stays in the customer's account while the software company runs the application plane, letting the vendor update the model while data never leaves
Scope: 'probably'
29:40 20VC: Who Wins in AI; Startup vs Incumbent, Infrastructure vs Application Layer, Bundled vs Unbundled Providers | From 150 LP Meetings to Closing $230M for Fund I; The Fundraising Process, What Worked, What Didn't and Lessons Learned with Tomasz Tunguz
Emad Mostaque · May 17, 2023
Open, auditable models are workable for sensitive healthcare data because you don't need to open all the underlying data — an on-device model can share only the minimum necessary fact with larger global models
Modern models are few-shot learners, so this isn't a classical big data problem; a small on-device model can hold private detail and pass only what's needed (e.g. that someone is old enough to drink, not their full record)
Scope: assumes fully auditable open models with no web-scraped data; architecture of big global models plus on-device models
11:36 20VC: Why the AI Bubble Will Be Bigger Than The Dot Com Bubble, Why AI Will Have a Bigger Impact Than COVID, Why No Models Used Today Will Be Used in a Year, Why All Models are Biased and How AI Kills Traditional Media with Emad Mostaque, Founder & CEO @
Douwe Kiela · Jun 30, 2023
Enterprise AI requires a clean separation between the data plane and the model plane, achieved by decoupling retrieval from generation so data stays in the customer's VPC while the model sits elsewhere.
Putting the model inside the customer's VPC gives full data privacy but no control, feedback or learning; a hybrid preserves privacy while retaining control.
30:36 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI
Enterprises need process driven synthesis not retrieval of existing facts
George Sivulka · Jan 22, 2025
What enterprise users actually want is a system that takes everything that already exists and answers process-driven questions about it (customer concentration, management strength, firm-specific investing criteria), not one that finds something that already exists
Judgments like 'is this a good investment' are the output of a process and are never explicitly stated in the source documents
22:05 20VC: Why All AI Companies Are Under-Valued | The Future of Foundation Models: Scaling Laws, Generalised vs Specialised, Commoditised? | From Unable to Afford Rent to Raising $130M From Index and Peter Thiel with George Sivulka @ Hebbia
George Sivulka · Jan 22, 2025
Roughly 90% of the questions enterprise users actually ask AI systems over their documents cannot be answered by searching the documents; they require work to be performed on top of the documents before an answer exists
Reviewing real customer queries at large finance firms showed people asking things like 'read all the documents and tell me every mention of AI' or 'what is our exposure to Silicon Valley Bank', not 'find me the quote'
Scope: based on observed query logs from large finance-firm deployments
23:28 20VC: Why All AI Companies Are Under-Valued | The Future of Foundation Models: Scaling Laws, Generalised vs Specialised, Commoditised? | From Unable to Afford Rent to Raising $130M From Index and Peter Thiel with George Sivulka @ Hebbia
Also on the record
Steeve Morin · Feb 24, 2025
RAG is a clever but dirty trick, because you are limited by how much data you can insert and therefore face an unsolved chunking problem
The injected window is fixed in size, so you must decide how to chunk the data
57:12 Rag is limited by unsolved chunking and context window constraints
George Sivulka · Jan 22, 2025
Large language models are good at thinking if given the right context, but they hardly ever have the right context — which is the problem RAG was the first real attempt to solve
Models answer from memory unless you retrieve and supply the data needed to answer correctly
20:28 Models need the right context rag is the first attempt to supply it
Your assistant can query this graph directly — 10 positions here, 19,646 across the corpus. Add 996.fm over MCP.