Skip to content

Debates

Is proprietary data still a critical competitive moat for building AI-powered products?

4 recorded positions from 3 people, first said Jun 19, 2023. The readings below are distinctions drawn on the record, not opposing camps.

Proprietary data matters less now small data suffices to bootstrap via fine tuning

Alex Lebrun · Jun 19, 2023 · hedged

Proprietary data is less and less of a competitive advantage; it mattered greatly in the last cycle but now you only need a little to bootstrap

With new pre-trained models, fine tuning is very efficient with a small amount of data — Nabla had to pay doctors to build 30,000 medical consultations only to bootstrap the product

Scope: you still need some data to bootstrap a model

15:03 20VC: Why No Models Today Will Be Used in a Year, Why Open Will Always Beat Closed in AI, Why Proprietary Data is Less Important Than Ever And Why EU AI Regulation is a Disaster with Alex Lebrun, Founder & CEO @ Nabla

Alex Lebrun · Jun 19, 2023

Because the heavy lifting happens in pretraining, a small amount of high-quality data is enough to make a model do what you need

Pretraining absorbs huge amounts of text, so fine tuning only needs small, high-quality datasets to steer behaviour

15:47 20VC: Why No Models Today Will Be Used in a Year, Why Open Will Always Beat Closed in AI, Why Proprietary Data is Less Important Than Ever And Why EU AI Regulation is a Disaster with Alex Lebrun, Founder & CEO @ Nabla

Akin Babayigit · Jun 26, 2023

Lack of proprietary data should not be the deciding factor when investing in new AI tool companies

You can build a hugely valuable business by offering tools and services that change people's lives without proprietary data, and the business model will change anyway — he already uses a powerful tool with little proprietary data

Scope: about the state of these companies today

51:11 20VC: Eight Pieces of Startup Advice that are BS: Why You Do Not Have to Love Your Space, It Is Ok To Do It For The Money, Focus Is Not Everything, Speed Is Not The Most Important Thing with Akin Babayigit, Co-Founder @ Tripledot Studios

Douwe Kiela · Jun 30, 2023

Proprietary data is essential for a deep tech AI-first startup but barely needed for a startup building on top of large language models

Deep tech AI companies need a big data flywheel where data is the moat, but LLMs are so sample efficient that you can do impressive things with very little data — possibilities that didn't exist a couple of years ago

Scope: depends on whether you are building the model or building on top of it

15:09 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI

Your assistant can query this graph directly — 4 positions here, 19,646 across the corpus. Add 996.fm over MCP.