Is proprietary data still a critical competitive moat for building AI-powered products?
4 recorded positions from 3 people, first said Jun 19, 2023. The readings below are distinctions drawn on the record, not opposing camps.
Proprietary data matters less now small data suffices to bootstrap via fine tuning
Alex Lebrun · Jun 19, 2023 · hedged
Proprietary data is less and less of a competitive advantage; it mattered greatly in the last cycle but now you only need a little to bootstrap
With new pre-trained models, fine tuning is very efficient with a small amount of data — Nabla had to pay doctors to build 30,000 medical consultations only to bootstrap the product
Scope: you still need some data to bootstrap a model
15:03 20VC: Why No Models Today Will Be Used in a Year, Why Open Will Always Beat Closed in AI, Why Proprietary Data is Less Important Than Ever And Why EU AI Regulation is a Disaster with Alex Lebrun, Founder & CEO @ Nabla
Alex Lebrun · Jun 19, 2023
Because the heavy lifting happens in pretraining, a small amount of high-quality data is enough to make a model do what you need
Pretraining absorbs huge amounts of text, so fine tuning only needs small, high-quality datasets to steer behaviour
15:47 20VC: Why No Models Today Will Be Used in a Year, Why Open Will Always Beat Closed in AI, Why Proprietary Data is Less Important Than Ever And Why EU AI Regulation is a Disaster with Alex Lebrun, Founder & CEO @ Nabla
Akin Babayigit · Jun 26, 2023
Lack of proprietary data should not be the deciding factor when investing in new AI tool companies
You can build a hugely valuable business by offering tools and services that change people's lives without proprietary data, and the business model will change anyway — he already uses a powerful tool with little proprietary data
Scope: about the state of these companies today
51:11 20VC: Eight Pieces of Startup Advice that are BS: Why You Do Not Have to Love Your Space, It Is Ok To Do It For The Money, Focus Is Not Everything, Speed Is Not The Most Important Thing with Akin Babayigit, Co-Founder @ Tripledot Studios
Douwe Kiela · Jun 30, 2023
Proprietary data is essential for a deep tech AI-first startup but barely needed for a startup building on top of large language models
Deep tech AI companies need a big data flywheel where data is the moat, but LLMs are so sample efficient that you can do impressive things with very little data — possibilities that didn't exist a couple of years ago
Scope: depends on whether you are building the model or building on top of it
15:09 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI
Your assistant can query this graph directly — 4 positions here, 19,646 across the corpus. Add 996.fm over MCP.