Skip to content

Debates

Do academic benchmarks measure what matters for AI model progress?

19 recorded positions from 10 people, first said Jun 30, 2023. They do not agree — the readings below are what each one actually argued.

Real workflow evals must replace academic benchmarks

Harry Stebbings · Apr 11, 2025

According to Glean's CPO/president (as relayed by the interviewer), Glean's internal evals are very different from public benchmark evals

41:40 20Product: How Scale AI and Harvey Build Product | Why PMs Are Wrong: They are not the CEOs of the Product | How to do Pre and Post Mortems Effectively and How to Nail PRDs | The Future of Product Management in a World of AI with Aatish Nayak

Edwin Chen · Jul 21, 2025

Benchmark-topping models are optimized for narrowly scoped academic homework-style problems, which is not the same as being good at the problems people actually face

It's the equivalent of making a model good at SAT problems rather than real-world tasks

45:43 20VC: Scaling to $1BN+ in Revenue with No Funding: Surge AI | The Most Insane Scaling Story in Tech |

Nick Frosst · Sep 1, 2025

What matters is whether a customer's task works as easily as possible, not benchmark scores or industry discourse

None of what they optimize for is reflected in the benchmarks that cycle through each year

17:21 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst

Nick Frosst · Sep 1, 2025

Popular AI benchmarks don't reflect the utility of models for real customers

Benchmarks like AIME test math reasoning that almost no workplace needs, and ARC-AGI is a pixel manipulation task no customer has ever asked a model to do; earlier benchmarks like LM1B and HellaSwag have already been abandoned

Scope: some benchmarks still contain good scientific work; evaluating emergent capabilities is interesting; stops short of calling them all bullshit

18:18 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst

Harry Stebbings · Sep 15, 2025

Current model evaluation methods like Humanity's Last Exam don't determine practical usage in society.

25:43 20VC: Mercor: From $1M to $500M in 17 Months: The Fastest Growing Company in the World | How to Think About Margins and Revenue Sustainability in AI | Why Evaluation Benchmarks in AI are BS Today with Brendan Foody

Brendan Foody · Sep 15, 2025

One of the largest inefficiencies in AI research is that prevailing benchmarks — Humanity's Last Exam, PhD-level reasoning, Olympiad math — are wholly disconnected from the outcomes consumers and enterprises actually care about.

What buyers want is a model that can build a financial model like a Goldman banker, produce consulting decks, or build a web app like a real engineer.

25:56 20VC: Mercor: From $1M to $500M in 17 Months: The Fastest Growing Company in the World | How to Think About Margins and Revenue Sustainability in AI | Why Evaluation Benchmarks in AI are BS Today with Brendan Foody

Brendan Foody · Sep 15, 2025

The MIT finding of widespread AI pilot failure confirms that benchmark achievements like Olympiad gold or PhD-level reasoning don't translate into enterprise usefulness

Evals are how you measure truth about what models can actually do, and the academic ones are disconnected from enterprise outcomes

34:53 20VC: Mercor: From $1M to $500M in 17 Months: The Fastest Growing Company in the World | How to Think About Margins and Revenue Sustainability in AI | Why Evaluation Benchmarks in AI are BS Today with Brendan Foody

Joelle Pineau · Nov 3, 2025

Selling AI into enterprises gives a far truer signal of what works than academic benchmarks, and that real-world feedback should guide which research ideas to pursue

Academic benchmarks give some signal but not the same as getting a model to do productive work; enterprise deployment yields new data and insights that narrow a huge space of research ideas

14:43 20VC: Cohere's Chief AI Officer on Why Scaling Laws Will Continue | Whether You Can Buy Success in AI with Talent Acquisitions | The Future of Synthetic Data & What It Means for Models | Why AI Coding is Akin to Image Generation in 2015 with Joelle Pineau

Brendan Foody · Jun 1, 2026

Academic benchmarks were disconnected from the outcomes enterprises care about, and pushing the frontier of evaluation toward end-to-end real workflows is now a critical research problem for the next generation of models

Benchmarks like GPQA, IMO and Humanity's Last Exam test academic problems no one really cares about, whereas what matters is whether a model can run an end-to-end workflow like building a financial model, a slide deck, or an entire SaaS application

48:16 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility, The Model is the Product | Token Spend Will Exceed Headcount Spend in 5 Years | The True Cost of Hiring AI Researchers in the Valley Today with Brendan Foody

Leaderboards are unhelpful except as consumer novelty

Ethan Mollick · Jul 31, 2024

The AI enthusiast community's obsession with weekly leaderboard changes is largely unnecessary; what matters is when models top out and how long that takes, not who leads at any given moment

All the big labs will keep building regardless, and normal users mostly just keep using ChatGPT rather than switching with each release; social media rewards buzz

Scope: fine to follow along if you're an enthusiast

7:36 20VC: Is More Compute the Answer to Model Performance | Why OpenAI Abandons Products, The Biggest Opportunities They Have Not Taken & Analysing Their Race for AGI | What Companies, AI Labs and Startups Get Wrong About AI with Ethan Mollick

Nick Frosst · Sep 1, 2025

Model leaderboards aren't very helpful

Scope: they are fun and fine in the consumer space where people want to try the newest thing

19:16 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst

Also on the record

Joelle Pineau · Nov 3, 2025

Evals are genuinely useful as unit tests that signal how a system performs on a particular dimension, but should not be treated as the ultimate goal or optimized against

Benchmarks give a directional signal and can be predictive of behaviour elsewhere, but the choice of evaluation depends on what kind of model and system you are building

45:01 Benchmarks are unit tests for knowledge not optimization targets

Edwin Chen · Jul 21, 2025

The right North Star metric for the industry would be whether models are progressing in fundamental capability rather than climbing meaningless leaderboards, and the best available proxy today is the variety and creativity of projects being run

Researchers should be able to pursue new ideas without being blocked by data, so more complex, diverse and creative projects approximate genuine capability progress

51:44 Project variety and creativity is the best proxy for real capability progress

Douwe Kiela · Jun 30, 2023

We no longer know how to evaluate the quality of language models, and using GPT-4 to evaluate other language models is the wrong approach

The field has resorted to GPT-4-as-judge, and data contamination means models are often trained on what they're evaluated on

18:56 Gpt4 as judge evaluation is flawed due to data contamination

Douwe Kiela · Jun 30, 2023

Static test-set benchmarks like Stanford HELM are the wrong way to measure model progress; evaluation should be dynamic, measuring how hard it is for an adversarial human to break the model

Adversarial attack success rate is a metric that keeps working over time — it used to be easy to make models fail and is getting harder, which tracks real progress

19:58 Dynamic adversarial attack success rate beats static benchmarks

Nick Frosst · Sep 1, 2025

Benchmark scores mostly reflect how much a model has been trained on those benchmarks, and can definitely be gamed

19:06 Benchmark scores reflect training on them and can be gamed

Aatish Nayak · Apr 11, 2025

Public legal benchmarks are the wrong instrument because they are multiple choice, whereas real legal work is open-ended and requires task-specific rubrics

Any lawyer will tell you there are a million options for what you could do; open-ended tasks like generating a chronology versus drafting a motion for summary judgment need their own rubrics

41:59 Open ended legal tasks need task specific rubrics not multiple choice benchmarks

Richard Socher · Apr 18, 2025

In AI, a claim to have the best model is quickly checked because people can play around with the model themselves, and they will call BS on false claims

Public playability is what makes a 'best model' claim credible

10:35 Public playability self corrects false best model claims

Alexander Embiricos · Feb 21, 2026

Benchmarks and evals deserve some weight as a measure of intelligence — especially meaningful progress on unsaturated evals — but must be paired with the vibes-based experience of actually using the model.

Whenever he talks to internal users or customers, he is surprised by how vibes-based the evaluation of what it feels like to work with a model is.

41:54 Unsaturated benchmarks plus vibes based use together

Your assistant can query this graph directly — 19 positions here, 19,646 across the corpus. Add 996.fm over MCP.