Is model hallucination a technical defect to eliminate or a trust and design problem to manage?
13 recorded positions from 12 people, first said May 17, 2023. They do not agree — the readings below are what each one actually argued.
Hallucination is a feature in creative contexts and a bug only where facts matter
Douwe Kiela · Jun 30, 2023
Hallucination is a spectrum rather than simply a feature or a bug — desirable for creative writing, unacceptable for enterprise-critical deployment
In creative use you want creativity and will revise the output anyway; in enterprise-critical situations you just want the model to do the right thing
Scope: depends on use case
10:00 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI
Cristóbal Valenzuela · Aug 28, 2023
Hallucination is a feature in creative image and video contexts and a bug only where facts are required
If you ask a language model for a fact you don't want a creative interpretation, but in a creative process you may want the model to ramble and be able to control the temperature
Scope: depends on the domain and whether facts are needed
18:54 20VC: Why AI Models are not a Moat, Where Does the Value in AI Accrue; Startups or Incumbents, What the World Has Got Wrong About AI, Why AI Needs a New Story and Who is the Right People to Tell it with Cris Valenzuela, Co-Founder & CEO @ Runway
Noam Shazeer · Aug 31, 2023
Hallucinations are a feature rather than a bug, and the right strategy is to launch something general and let the first use cases be ones where hallucination helps
If models hallucinate — which they certainly do and Character advertises — then the use cases that emerge first will naturally be ones where hallucination is an asset, like entertainment, emotional support and fun
Scope: productivity use cases are welcome too; he wants adoption to happen naturally based on what the technology is good at
22:43 20VC: Spending $2M to Train a Single AI Model: What Matters More; Model Size or Data Size | Hallucinations: Feature or Bug | Will Everyone Have an AI Friend in the Future & Raising $150M from a16z with Noam Shazeer, Co-Founder & CEO @ Character.ai
Christian Kleinerman · Sep 22, 2023
Creative industries are the sweetest spot for generative AI
What other industries call a bug or hallucination is a feature in creative work — producing something novel or a mix of existing things is a goodness there, whereas other verticals require correctness and data maturity
Scope: other verticals will follow, but need correctness and data maturity
9:04 20VC: Are Foundation Models Becoming Commoditised? Do OpenAI and Anthropic of the World Have a Sustaining Moat? Why Smaller Models May Work Better? Why Incumbents with Data Power Win the AI War with Christian Kleinerman, SVP Product @ Snowflake
Hallucination tolerance scales with application risk so agentic investment pays off
Scott Belsky · Dec 20, 2023
Hallucination is a feature rather than a bug in discovery and generative creative contexts, and is only a serious problem in mission-critical applications
For music discovery or generative fill imagining what's behind an object, an invented answer is the point and you can simply run it again; for things like drafting NDAs a hallucination can get you in trouble
Scope: depends on the application domain
8:23 20VC Roundtable: Spotify, Adobe & Linkedin CPOs on How AI Changes The Future of Product, Why AI is Now the Product, How TikTok Changed Product, Why Cost is the Biggest Barrier to LLM Usage & Why Incumbents Can Adopt AI Faster Than Any Prior Innovation Cyc
David Luan · Jun 24, 2024 · hedged
Chatbots and agents are speciating into different kinds of technology with different requirements, most visibly around hallucination — which is a feature for chatbots and image generators but disqualifying for agents.
Hallucination gives creative tools useful novelty for blank-page problems, but an agent doing your taxes or managing shipping containers must not invent things.
16:00 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept
Jonathan Ross · Feb 17, 2025
Money invested in agentic AI today will not be wasted, because whether hallucinations are disqualifying depends on how high-risk the industry is
Low-risk applications like Perplexity work fine today because users can check the links and it's entertainment-grade; positioning early for a wave pays off when it arrives, as Groq's seven years of waiting did
Scope: depends on risk level of the industry; requires being in the right position
74:30 20VC: NVIDIA vs Groq: The Future of Training vs Inference | Meta, Google, and Microsoft's Data Center Investments: Who Wins | Data, Compute, Models: The Core Bottlenecks in AI & Where Value Will Distribute with Jonathan Ross, Founder @ Groq
Trust is a design problem not a hallucination problem
Emad Mostaque · May 17, 2023
The biggest misconception about AI is hallucination: expecting full factual accuracy from a model with 10,000-to-50,000-to-one compression is wrong, and the fix is tying models into proper systems rather than using them one-on-one
The technology wasn't built for factual recall — that it does what it does at that compression ratio is miraculous; people should instead think about the data journey and provenance across embeddings and other systems
60:36 20VC: Why the AI Bubble Will Be Bigger Than The Dot Com Bubble, Why AI Will Have a Bigger Impact Than COVID, Why No Models Used Today Will Be Used in a Year, Why All Models are Biased and How AI Kills Traditional Media with Emad Mostaque, Founder & CEO @
Christian Lanng · Sep 27, 2023
Eliminating hallucinations is the wrong goal; trust is a design problem and the answer is showing why and when the model knows something
The more powerful the model, the more powerful the hallucinations; humans need designed cues to trust AI as a coworker, like a self-driving Tesla rendering the road purely to build trust, or a model asking a clarifying question even when it already knows the answer
Scope: in the work environment specifically
52:04 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng
Also on the record
Vlad Tenev · Jul 14, 2025
Hallucination remains AI's biggest shortfall, and a system that prevents hallucination by design — checking every step of reasoning against a clear measure of correctness — will change everything
Current models need scaffolding against hallucination; a verified-output approach guarantees each reasoning step follows from the last
9:23 Verified reasoning architecture can eliminate hallucination by design
Des Traynor · Nov 15, 2023
For customer support AI, trustworthiness and staying on topic matter more than answer coverage — it is worth prompting so conservatively that you occasionally lose a valid answer.
Customers need every answer Fin gives to be high confidence; a bank's support has no official policy on current events and neither should its bot.
14:37 Prioritize conservative high confidence answers over coverage in support ai
Alex Lebrun · Jun 19, 2023
Hallucination is inherent to LLM design rather than a defect, because the model cannot not produce an output
By construction the LLM has to output something, so if there is nothing to say it makes up something that looks natural
14:13 Hallucination is structurally inherent since the model must always produce an output
Alex Lebrun · Jun 19, 2023
Feeding an LLM curated, trusted data does not make its output trustworthy
An LLM is a probabilistic autocomplete that will complete a sentence at any cost — it can start with the beginning of one true sentence and finish with the end of another, producing perfectly-formed but factually wrong output; so removing noisy sources like Reddit from training data would not eliminate wrong answers
22:40 Curated trusted input data does not guarantee trustworthy output since llms blend sentences probabilistically
Your assistant can query this graph directly — 13 positions here, 19,646 across the corpus. Add 996.fm over MCP.