Do academic benchmarks measure what matters for AI model progress?
19 recorded positions from 10 people, first said Jun 30, 2023. They do not agree — the readings below are what each one actually argued.
Real workflow evals must replace academic benchmarks
Harry Stebbings · Apr 11, 2025
According to Glean's CPO/president (as relayed by the interviewer), Glean's internal evals are very different from public benchmark evals
41:40 20Product: How Scale AI and Harvey Build Product | Why PMs Are Wrong: They are not the CEOs of the Product | How to do Pre and Post Mortems Effectively and How to Nail PRDs | The Future of Product Management in a World of AI with Aatish Nayak
Edwin Chen · Jul 21, 2025
Benchmark-topping models are optimized for narrowly scoped academic homework-style problems, which is not the same as being good at the problems people actually face
It's the equivalent of making a model good at SAT problems rather than real-world tasks
45:43 20VC: Scaling to $1BN+ in Revenue with No Funding: Surge AI | The Most Insane Scaling Story in Tech |
Nick Frosst · Sep 1, 2025
What matters is whether a customer's task works as easily as possible, not benchmark scores or industry discourse
None of what they optimize for is reflected in the benchmarks that cycle through each year
17:21 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst
Nick Frosst · Sep 1, 2025
Popular AI benchmarks don't reflect the utility of models for real customers
Benchmarks like AIME test math reasoning that almost no workplace needs, and ARC-AGI is a pixel manipulation task no customer has ever asked a model to do; earlier benchmarks like LM1B and HellaSwag have already been abandoned
Scope: some benchmarks still contain good scientific work; evaluating emergent capabilities is interesting; stops short of calling them all bullshit
18:18 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst
Harry Stebbings · Sep 15, 2025
Current model evaluation methods like Humanity's Last Exam don't determine practical usage in society.
25:43 20VC: Mercor: From $1M to $500M in 17 Months: The Fastest Growing Company in the World | How to Think About Margins and Revenue Sustainability in AI | Why Evaluation Benchmarks in AI are BS Today with Brendan Foody
Brendan Foody · Sep 15, 2025
One of the largest inefficiencies in AI research is that prevailing benchmarks — Humanity's Last Exam, PhD-level reasoning, Olympiad math — are wholly disconnected from the outcomes consumers and enterprises actually care about.
What buyers want is a model that can build a financial model like a Goldman banker, produce consulting decks, or build a web app like a real engineer.
25:56 20VC: Mercor: From $1M to $500M in 17 Months: The Fastest Growing Company in the World | How to Think About Margins and Revenue Sustainability in AI | Why Evaluation Benchmarks in AI are BS Today with Brendan Foody
Brendan Foody · Sep 15, 2025
The MIT finding of widespread AI pilot failure confirms that benchmark achievements like Olympiad gold or PhD-level reasoning don't translate into enterprise usefulness
Evals are how you measure truth about what models can actually do, and the academic ones are disconnected from enterprise outcomes
34:53 20VC: Mercor: From $1M to $500M in 17 Months: The Fastest Growing Company in the World | How to Think About Margins and Revenue Sustainability in AI | Why Evaluation Benchmarks in AI are BS Today with Brendan Foody
Joelle Pineau · Nov 3, 2025
Selling AI into enterprises gives a far truer signal of what works than academic benchmarks, and that real-world feedback should guide which research ideas to pursue
Academic benchmarks give some signal but not the same as getting a model to do productive work; enterprise deployment yields new data and insights that narrow a huge space of research ideas
14:43 20VC: Cohere's Chief AI Officer on Why Scaling Laws Will Continue | Whether You Can Buy Success in AI with Talent Acquisitions | The Future of Synthetic Data & What It Means for Models | Why AI Coding is Akin to Image Generation in 2015 with Joelle Pineau
Brendan Foody · Jun 1, 2026
Academic benchmarks were disconnected from the outcomes enterprises care about, and pushing the frontier of evaluation toward end-to-end real workflows is now a critical research problem for the next generation of models
Benchmarks like GPQA, IMO and Humanity's Last Exam test academic problems no one really cares about, whereas what matters is whether a model can run an end-to-end workflow like building a financial model, a slide deck, or an entire SaaS application
48:16 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility, The Model is the Product | Token Spend Will Exceed Headcount Spend in 5 Years | The True Cost of Hiring AI Researchers in the Valley Today with Brendan Foody
Leaderboards are unhelpful except as consumer novelty
Ethan Mollick · Jul 31, 2024
The AI enthusiast community's obsession with weekly leaderboard changes is largely unnecessary; what matters is when models top out and how long that takes, not who leads at any given moment
All the big labs will keep building regardless, and normal users mostly just keep using ChatGPT rather than switching with each release; social media rewards buzz
Scope: fine to follow along if you're an enthusiast
7:36 20VC: Is More Compute the Answer to Model Performance | Why OpenAI Abandons Products, The Biggest Opportunities They Have Not Taken & Analysing Their Race for AGI | What Companies, AI Labs and Startups Get Wrong About AI with Ethan Mollick
Nick Frosst · Sep 1, 2025
Model leaderboards aren't very helpful
Scope: they are fun and fine in the consumer space where people want to try the newest thing
19:16 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst
Also on the record
Joelle Pineau · Nov 3, 2025
Evals are genuinely useful as unit tests that signal how a system performs on a particular dimension, but should not be treated as the ultimate goal or optimized against
Benchmarks give a directional signal and can be predictive of behaviour elsewhere, but the choice of evaluation depends on what kind of model and system you are building
45:01 Benchmarks are unit tests for knowledge not optimization targets
Edwin Chen · Jul 21, 2025
The right North Star metric for the industry would be whether models are progressing in fundamental capability rather than climbing meaningless leaderboards, and the best available proxy today is the variety and creativity of projects being run
Researchers should be able to pursue new ideas without being blocked by data, so more complex, diverse and creative projects approximate genuine capability progress
51:44 Project variety and creativity is the best proxy for real capability progress
Douwe Kiela · Jun 30, 2023
We no longer know how to evaluate the quality of language models, and using GPT-4 to evaluate other language models is the wrong approach
The field has resorted to GPT-4-as-judge, and data contamination means models are often trained on what they're evaluated on
18:56 Gpt4 as judge evaluation is flawed due to data contamination
Douwe Kiela · Jun 30, 2023
Static test-set benchmarks like Stanford HELM are the wrong way to measure model progress; evaluation should be dynamic, measuring how hard it is for an adversarial human to break the model
Adversarial attack success rate is a metric that keeps working over time — it used to be easy to make models fail and is getting harder, which tracks real progress
19:58 Dynamic adversarial attack success rate beats static benchmarks
Nick Frosst · Sep 1, 2025
Benchmark scores mostly reflect how much a model has been trained on those benchmarks, and can definitely be gamed
19:06 Benchmark scores reflect training on them and can be gamed
Aatish Nayak · Apr 11, 2025
Public legal benchmarks are the wrong instrument because they are multiple choice, whereas real legal work is open-ended and requires task-specific rubrics
Any lawyer will tell you there are a million options for what you could do; open-ended tasks like generating a chronology versus drafting a motion for summary judgment need their own rubrics
41:59 Open ended legal tasks need task specific rubrics not multiple choice benchmarks
Richard Socher · Apr 18, 2025
In AI, a claim to have the best model is quickly checked because people can play around with the model themselves, and they will call BS on false claims
Public playability is what makes a 'best model' claim credible
10:35 Public playability self corrects false best model claims
Alexander Embiricos · Feb 21, 2026
Benchmarks and evals deserve some weight as a measure of intelligence — especially meaningful progress on unsaturated evals — but must be paired with the vibes-based experience of actually using the model.
Whenever he talks to internal users or customers, he is surprised by how vibes-based the evaluation of what it feels like to work with a model is.
41:54 Unsaturated benchmarks plus vibes based use together
Your assistant can query this graph directly — 19 positions here, 19,646 across the corpus. Add 996.fm over MCP.