Skip to content

Debates

Is model hallucination a technical defect to eliminate or a trust and design problem to manage?

13 recorded positions from 12 people, first said May 17, 2023. They do not agree — the readings below are what each one actually argued.

Hallucination is a feature in creative contexts and a bug only where facts matter

Douwe Kiela · Jun 30, 2023

Hallucination is a spectrum rather than simply a feature or a bug — desirable for creative writing, unacceptable for enterprise-critical deployment

In creative use you want creativity and will revise the output anyway; in enterprise-critical situations you just want the model to do the right thing

Scope: depends on use case

10:00 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI

Cristóbal Valenzuela · Aug 28, 2023

Hallucination is a feature in creative image and video contexts and a bug only where facts are required

If you ask a language model for a fact you don't want a creative interpretation, but in a creative process you may want the model to ramble and be able to control the temperature

Scope: depends on the domain and whether facts are needed

18:54 20VC: Why AI Models are not a Moat, Where Does the Value in AI Accrue; Startups or Incumbents, What the World Has Got Wrong About AI, Why AI Needs a New Story and Who is the Right People to Tell it with Cris Valenzuela, Co-Founder & CEO @ Runway

Noam Shazeer · Aug 31, 2023

Hallucinations are a feature rather than a bug, and the right strategy is to launch something general and let the first use cases be ones where hallucination helps

If models hallucinate — which they certainly do and Character advertises — then the use cases that emerge first will naturally be ones where hallucination is an asset, like entertainment, emotional support and fun

Scope: productivity use cases are welcome too; he wants adoption to happen naturally based on what the technology is good at

22:43 20VC: Spending $2M to Train a Single AI Model: What Matters More; Model Size or Data Size | Hallucinations: Feature or Bug | Will Everyone Have an AI Friend in the Future & Raising $150M from a16z with Noam Shazeer, Co-Founder & CEO @ Character.ai

Christian Kleinerman · Sep 22, 2023

Creative industries are the sweetest spot for generative AI

What other industries call a bug or hallucination is a feature in creative work — producing something novel or a mix of existing things is a goodness there, whereas other verticals require correctness and data maturity

Scope: other verticals will follow, but need correctness and data maturity

9:04 20VC: Are Foundation Models Becoming Commoditised? Do OpenAI and Anthropic of the World Have a Sustaining Moat? Why Smaller Models May Work Better? Why Incumbents with Data Power Win the AI War with Christian Kleinerman, SVP Product @ Snowflake

Hallucination tolerance scales with application risk so agentic investment pays off

Scott Belsky · Dec 20, 2023

Hallucination is a feature rather than a bug in discovery and generative creative contexts, and is only a serious problem in mission-critical applications

For music discovery or generative fill imagining what's behind an object, an invented answer is the point and you can simply run it again; for things like drafting NDAs a hallucination can get you in trouble

Scope: depends on the application domain

8:23 20VC Roundtable: Spotify, Adobe & Linkedin CPOs on How AI Changes The Future of Product, Why AI is Now the Product, How TikTok Changed Product, Why Cost is the Biggest Barrier to LLM Usage & Why Incumbents Can Adopt AI Faster Than Any Prior Innovation Cyc

David Luan · Jun 24, 2024 · hedged

Chatbots and agents are speciating into different kinds of technology with different requirements, most visibly around hallucination — which is a feature for chatbots and image generators but disqualifying for agents.

Hallucination gives creative tools useful novelty for blank-page problems, but an agent doing your taxes or managing shipping containers must not invent things.

16:00 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept

Jonathan Ross · Feb 17, 2025

Money invested in agentic AI today will not be wasted, because whether hallucinations are disqualifying depends on how high-risk the industry is

Low-risk applications like Perplexity work fine today because users can check the links and it's entertainment-grade; positioning early for a wave pays off when it arrives, as Groq's seven years of waiting did

Scope: depends on risk level of the industry; requires being in the right position

74:30 20VC: NVIDIA vs Groq: The Future of Training vs Inference | Meta, Google, and Microsoft's Data Center Investments: Who Wins | Data, Compute, Models: The Core Bottlenecks in AI & Where Value Will Distribute with Jonathan Ross, Founder @ Groq

Trust is a design problem not a hallucination problem

Emad Mostaque · May 17, 2023

The biggest misconception about AI is hallucination: expecting full factual accuracy from a model with 10,000-to-50,000-to-one compression is wrong, and the fix is tying models into proper systems rather than using them one-on-one

The technology wasn't built for factual recall — that it does what it does at that compression ratio is miraculous; people should instead think about the data journey and provenance across embeddings and other systems

60:36 20VC: Why the AI Bubble Will Be Bigger Than The Dot Com Bubble, Why AI Will Have a Bigger Impact Than COVID, Why No Models Used Today Will Be Used in a Year, Why All Models are Biased and How AI Kills Traditional Media with Emad Mostaque, Founder & CEO @

Christian Lanng · Sep 27, 2023

Eliminating hallucinations is the wrong goal; trust is a design problem and the answer is showing why and when the model knows something

The more powerful the model, the more powerful the hallucinations; humans need designed cues to trust AI as a coworker, like a self-driving Tesla rendering the road purely to build trust, or a model asking a clarifying question even when it already knows the answer

Scope: in the work environment specifically

52:04 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng

Also on the record

Vlad Tenev · Jul 14, 2025

Hallucination remains AI's biggest shortfall, and a system that prevents hallucination by design — checking every step of reasoning against a clear measure of correctness — will change everything

Current models need scaffolding against hallucination; a verified-output approach guarantees each reasoning step follows from the last

9:23 Verified reasoning architecture can eliminate hallucination by design

Des Traynor · Nov 15, 2023

For customer support AI, trustworthiness and staying on topic matter more than answer coverage — it is worth prompting so conservatively that you occasionally lose a valid answer.

Customers need every answer Fin gives to be high confidence; a bank's support has no official policy on current events and neither should its bot.

14:37 Prioritize conservative high confidence answers over coverage in support ai

Alex Lebrun · Jun 19, 2023

Hallucination is inherent to LLM design rather than a defect, because the model cannot not produce an output

By construction the LLM has to output something, so if there is nothing to say it makes up something that looks natural

14:13 Hallucination is structurally inherent since the model must always produce an output

Alex Lebrun · Jun 19, 2023

Feeding an LLM curated, trusted data does not make its output trustworthy

An LLM is a probabilistic autocomplete that will complete a sentence at any cost — it can start with the beginning of one true sentence and finish with the end of another, producing perfectly-formed but factually wrong output; so removing noisy sources like Reddit from training data would not eliminate wrong answers

22:40 Curated trusted input data does not guarantee trustworthy output since llms blend sentences probabilistically

Your assistant can query this graph directly — 13 positions here, 19,646 across the corpus. Add 996.fm over MCP.