Cold open
Data size matters even more than model size. If you train a smaller model on more data for longer, then you get a better model. There was this Google memo from an internal Google employee who who had written that OpenAI and Google have no moat. For me as an AI researcher, when I read that memo, I was like, this person has no idea what they’re talking about.
Intro
Welcome back to 20 with me, Harry Stebbings. And today, we delve further into the world of AI with expert who’s been in the industry for the last fifteen years, Douwe Kiela, cofounder and CEO of Contextual AI, building the contextual language model to power the future of businesses. Last month, Contextual closed a $20,000,000 funding round, including Capital, Sarah Guo, Elad Gil, and twenty VC. He’s also an adjunct professor in Symbolic Systems at Stanford University. Previously, he was the head of research at Hugging Face, and before that, a research scientist at Facebook AI Research.
But before we dive into the show state,
· Sponsor read0 min · 485 words
did you know that over 50% of your day is filled with tedious tasks? What would you do if you got half a day back? Well, you can with Coda, the all in one platform that changes the way your team works together. And Coda just introduced an AI powered work assistant to take the busy out of work. With Coda, your team’s important workflows and content already live in one place, and Coda AI helps each team member focus on the highest priority work even as priority shift across the team.
By taking over recurring, and let’s be honest, very tedious work often, Coda AI empowers team members to prioritize longer term and strategic work to accomplish goals. Coda AI not only makes space for collaboration in each person’s workday, but can make it easier to stay in the loop and share informed opinions. So if you want to work a system that lets you get back to work, you can get started with Coda AI today for free. Head over to coda.io/20vc, that’s coda.io, and get started for free, coda.io/20bc.
And speaking of tools we cannot live without like coda, we have to talk about Brex, the all in one financial stack trusted by founders. Founders have to think globally in order to open new markets, unlock cost savings, and gain access to new talent. That’s why having the right financial stack is more important than ever, and that’s where Brex comes in. With Brex, you get a high limit corporate card, a high yield business account with up to $6,000,000 in FDIC protection, and bill pay, all built with a global first mindset.
Brexit enables you to operate in more countries and currencies than any other provider. So you can pay vendors, run payroll, and make international payments faster. We all know that’s a must have for startups at any stage of growth. So are you ready to learn more about Brexit’s global first solution? Visit brex.com/20vc. That’s brex.com/20vc. And finally, AngelList is fast becoming the center of the venture ecosystem. The startup world is just buzzing about how fast they’ve been shipping products that meaningfully improve the lives of both startups and fund managers and their investors.
For startups, AngelList reduces friction of cap table management, banking, and fundraising all in one place. Teams can focus on scaling and let AngelList handle the rest. Thousands of startups have moved their cap tables to AngelList in the past year. AngelList also supports large venture funds and their teams with an automated software first approach and the best customer service in the industry. Fund managers can focus on making great deals while AngelList handles reporting, taxes, compliance, and more. So if you’re ready to scale your startup or fund with the platform at the center of it all, visit angellist.com/20vc to get started.
One day. Hello. You have now arrived at your destination.
Conversation
Douwe, I am excited for this. Listen, we’ve chatted before. We’ve known each other for a while. But thank you so much for joining me today.
Yeah. Thanks very much for having me on the show. I’m a big fan.
Oh, it’s very, very kind of you. I literally paid you $25,000 to say that. But my question to you is, it’s such a hot space and there’s very few people who’ve actually been in it for a while. You are one of them. How did you first make your way into the world of ML and NLP first?
My journey has been a little bit unusual, actually. So when I was in high school in The Netherlands, I wanted to be a cool kid during the day, but at night, I was secretly fascinated by computers. So I taught myself to code. So then by the time I had to go to college and go study something, I thought I already knew everything about computer science. So I decided to study philosophy instead, radical departure from what I had been interested in at the time, but it was fascinating.
I use it still every day, I think. But then at some point in my career, it became clear that I had to start making money, so I needed a real job, and philosophy is not really a real job. And I did some logic in between foundations of math, which is also not really a real job. So I decided to study computer science after all. So I went to Cambridge in The UK, and so that’s really where I started doing NLP, natural language processing. One of my internships was done at Microsoft Research in New York with a very famous researcher called Leon Botu, who is Yann LeCun’s kind of, I wouldn’t say sidekick because that doesn’t really do justice to what he’s done.
So one of the godfathers of deep learning. And I had the opportunity to work with him. That was really an amazing time. So afterwards, when Yann and Leon started FAIR, Facebook AI Research, I joined that out of my PhD and that really kicked off my career.
I actually spoke to Leon as part of the prep for my interview with Yann. Amazing, amazing person. But I do wanna talk about that five years that you spent at Facebook’s Research team. It’s such a transformational team, as you said, with incredible individuals like Yann and Leon. What are your biggest takeaways from that experience, and how did it impact how you think today?
I learned so much there, mostly around how to focus your research direction. So I think initially I was doing all kinds of weird stuff, and it took me a while to figure out that a very clear real world application for the research that you’re doing makes it much more valuable than than going off on a tangent and maybe being a bit too far ahead of the rest of the field. A special place, and I think actually they don’t get enough credit for the impact that they’ve had on the world.
Facebook or Meta in general, actually. So almost every web app in the world runs on React, which is an open source project coming out of Meta.
Where does Contextual come from? What was that I’ve gotta do this. This is now the idea and the time.
Yeah. So this really started at the beginning of the year. We just saw this great need after ChatGPT had gone viral. We saw this great excitement in the world, at the same time, a lot of disappointment about it not being quite ready yet for real world adaption in enterprises where you actually want to use this technology. So we decided that now really is the right time to build a company to try to tackle that, And we think it’s still very early innings in the game, so I think a lot of people sometimes think that the game has been played, it’s just getting started.
Okay. So you said there about not being ready for like traditional adoption. I think a lot of the general public would say absolutely, it’s cool. What makes it not ready for general adoption do you think?
So there are a couple of just really big issues. Hallucination, these models make things up with very high confidence. Attribution, we don’t know why they’re saying what they’re saying. We can’t really trace it back to anything. There’s compliance issues, so we can’t really remove information from them. Tricky from a GDPR perspective, for example. We can’t revise information. We can’t keep it up to date. There’s massive data privacy issues where you have to send your very valuable company data. If you’re an enterprise, you have to send that to somebody else’s servers.
These models are also quite inefficient still, so you can make them much faster. What we are building at Contextual is a different kind of language model. We’re really thinking about this as the next generation of language models where we think about it from first principles for enterprise use cases. And what that means is that we wanna solve all of these problems by being a bit smarter about the architecture. And the architecture we’re specifically basing it on is retrieval augmented generation, which is something that me and my colleagues at FAIR came up with in 2020.
And what you do there is you decouple the memory from the generative capacity of the large language model, and this allows you to ground the generations from the language model in the things you retrieved in your memory essentially. So you get much less hallucination, you get attribution for free, you can always update the memory so you can remove information on the fly, you can add things, you can revise them, you can have a stream of memory. It’s much more efficient because you’re compressing a lot of the compute inside the memory, and you have a very clean separation between the data plane and the model plane, as they call it, which means that you can have better data privacy guarantees.
I have so many questions for you. First off, you said that about kind of attribution of knowing where it comes from and being a bit of a black box. Had someone on the show the other day and they said, open or closed, you still don’t really know what’s going on in the core foundational model layer. The open or closed is not really the point. You still don’t know. Is that true? And how do we think about actual true transparency of knowing what’s going on and why it’s producing what it is?
So we’re not gonna be able to know why a neural network does what it does at the scale that neural networks operate. So this is kind of like your own brain. Right? Like, I think your behavior is relatively predictable, so that goes for every human. Right? We all, like, can predict each other’s actions, but I have no idea what’s going on in your brain, and I will never know. There’s no way I can know. The only way I can kinda find out is by asking you.
But if you train the architecture the way we are training it right now, then at least you make sure that the model has learned to rely on the information that it finds. And that gives you much stronger attribution than if it’s just predicting the next word based on what it has seen before. Because it doesn’t have this ability from birth, basically, to find relevant information and ground its generation on that thing it found.
Another thing that I have to ask, you mentioned the word hallucinations there. I had em out at Stability on the show, and he said hallucinations are a feature, not a bug, which I thought was a very tweetable statement. Do you agree hallucinations are a feature, not a bug?
It’s a great quote, but as always, it’s a bit more nuanced than that. I think in some cases, is a feature if you want to use a language model for creative writing. And if you want it to be really, really creative, then you probably want it to hallucinate. So in a way, it’s a spectrum of groundedness and hallucination. If you really care about the language model doing the right thing and you want to deploy it in an enterprise critical situation, then you really don’t want it to be creative, you don’t want it to hallucinate, just want it to do what it has to do.
But if you want to use it for a creative writing exercise, then sure, you can have it hallucinate because you’re gonna revise whatever it gives you anyway.
He also mentioned the multi model aspect. Another guest on the show said that the winners in startup land will be determined by those who can bluntly switch models faster than anyone else, and that today’s models will be unusable in a year. Do you agree with them?
At this particular point in time, probably, yes. But just because the field is moving so incredibly quickly. And so I think that in the next year we’re gonna see lots of other models coming out and if you can have a language model agnostic AI company that relies on language models, then that would give you a competitive advantage. But at the same time, that also runs the risk of you relying on other people’s language models. It’s a bit of a tricky situation.
Do you think there will be many more startup language model companies?
This technology is really going to change the world and every aspect of it. So there are definitely a couple of incumbents, but they are also focused on very specific parts of the market. If you look at Anthropic and OpenAI, I think they’re really chasing for this idea of AGI and they’re relatively consumer facing. If you focus less on this idea of artificial general intelligence and you want to have something a bit more like artificial specialized intelligence where you just need the model to do what it needs to do, you don’t need it to know about Shakespeare or quantum mechanics or things that.
You just need it to solve your business’ problem. In that case, I I think there’s still a lot of room for innovation, and that’s where we are trying to innovate.
We’ve seen size of model matter less and less, it would seem. How do you think about the importance of size of model today? And does it matter as much as it used to? And will it matter even less with every year, month, day?
Yeah. Great question. I think Sam Altman had this interesting quote where he was saying that he thought models would stop growing in size. GPT-four kind of hit this ceiling. I think that’s probably right, but not really because size doesn’t matter. It’s just that data size matters even more than model size. And I think the Llama paper out of Meta really brilliantly showed this. If you train a smaller model on more data for longer, then you get a better model. So you get more bang for your buck if you train it on more data rather than having more parameters.
But in an ideal world, if you had infinite compute budget and infinite data, then you would train the biggest possible model because that’s the most likely to give you emergent capabilities as we call them in the field.
Does that take more time then? If you have smaller models with more data, given you need to feed more data through the model, does it not take more time than if you needed less data going through the model?
It depends. So it’s a trade off here. Right? But these big models also need a lot of data. So it really is a function of the number of GPUs that you have available. Let’s say you have a thousand GPUs, you can choose to train a huge model on relatively little data, and it will be okay, but it will be undertrained. So you have some sort of optimal point where you can train the model to perfection. In the field, we were underestimating where that optimal point is.
And it seems that data is much more important than model size when it comes to what’s optimal.
If we take this to a next, like, layer deeper, data more important than model size. What does that mean then in terms of who’s advantaged? Does that mean that startups are more advantaged actually than incumbents? I I don’t understand. Who’s more advantaged in that case? Is it incumbents because they already have existing massive data moats?
It depends on where the data comes from. I think incumbents definitely have an advantage there, but only some of them. And a lot of the data is just freely available on the Internet. Right? So the Llama model was not trained on any proprietary data. It was just trained on open data on the web. And there’s a lot more data to be had there. And as a society, we’re generating a ton of data every day to add to that big pile of data. So you can really train very high quality language models just on public data on the web.
But I think if you look at the secret sauce to a lot of these other models, like why is GPT-four so awesome, a part of that is that they went through enormous lengths to get like special data that nobody else has. So allegedly, they did this whisperer project where they’re very good at transcribing audio because that would allow them to transcribe, like, all of the podcasts in the world, which gives you very high quality language. If you can train on that language but nobody else has it, that puts you in a position of advantage.
How important is proprietary data? The main reason I would say why VCs are turning down startup AI companies is because they do not have a proprietary dataset to operate against, and they are defined as like a thin layer of generative AI on top of a foundational model. How important is proprietary data, do you think, for startups innovating in the space?
If you want to build a deep tech AI startup, then you really want to get a big data flywheel going. You want to start with a lot of data and then have a way to generate lots more data, and that data is gonna be your moat. But I think one of the interesting things about these large language models is that they’re incredibly sample efficient or data efficient. So you can do cool things with them with relatively little data that just previously just wasn’t possible. That unlocks all kinds of possibilities that just didn’t exist even a couple of years ago.
So on the one hand, yes, you need lots of data if you wanna build like big AI first things. But at the same time, if you want to do a startup that builds on top of this technology, you need very little data to get started. Bit of a tangent. But one of the use cases I’ve been seeing now for GPT-four is actually that people are using it to generate data and then they’re training on that data with cheaper models. So g p d four might end up disrupting not knowledge workers necessarily, but it might just disrupt like Mechanical Turk and is just an annotator on steroids.
And you can use all of that data to get much more custom models that you can then deploy very cheaply on specialized use cases. That’s a quite interesting development.
Pre trained data changes a lot. Can you just help anyone who doesn’t know understand, what is pre trained data? How does it change the game for a lot of companies that don’t have existing data modes?
Maybe it’s useful to kind of go through the steps. If you want to build your own ChatGPT, like, do you need? And so the first thing you need is a core pre trained model, and this tends to be just trained on the web. The task you’re training it on is just next word prediction. Then once you have that core model, then you want to do supervised fine tuning. So essentially, you want to fix the user interface to that model because the model doesn’t really know how to follow instructions, for example.
So you want the model to listen to you, but it has only been trained on predicting the next word. So it doesn’t really know how to do that. So that supervised fine tuning, that’s also proprietary data. You can get a much better model out of that. And then the final step is RLHF, reinforcement learning from human feedback, where you get this feedback loop to make the model even better for your specific use case even if you don’t have signal at the word level. You just have signal at the sequence level.
So you can tell it like, okay, that was a good response or that wasn’t a good response, but you can’t tell it like, what did you do wrong necessarily. If you go through those three steps, then you get ChatGPT. It’s as easy as that.
Which company do you think has the best data acquisition flywheel? When you look at them today, who do you admire and respect most?
OpenAI. They haven’t even really trained as far as I know on the data that comes out of ChatGPT going viral. And so they had ChatGPT. It went viral. This led to this giant giant data moat that they haven’t even really used yet. So in terms of data moats, and and maybe you you’ve seen this come by actually, there was this Google memo from an internal Google employee who who had written that OpenAI and Google have no moat. And for me as an AI researcher, when I read that memo, I was like, this person has no idea what they’re talking about.
Why are they
role
and don’t understand?
These places have a giant moat because as I said, it’s really all about data. And OpenAI has this very deep understanding of how people want to use language models, basically nobody else has. And they have this giant economy of scale where they can serve up language models very cheaply because they get so many requests coming in at the same time. So they have a giant moat. So I’m a big fan of open source. Right? I would like it to be true that with open source, we could just keep up with all of that.
But I think that’s just incredibly naive.
What do you think are the biggest challenges that they face? Because I think we all dismissed Google quite significantly, if I’m honest. And then Bard came out and was pretty impressive. How do you evaluate Bard in Google’s display actually?
Language model evaluation is a whole separate topic. It’s a super interesting question actually. I’ve been fascinated by AI evaluation for a really long time. And the answer is we don’t really know how to evaluate the quality of these models anymore. So what we’ve seen people do in the field now is they’re using GPT-four to evaluate the quality of other language models. That just feels wrong. So I think there’s a giant opportunity in the market actually for a startup or several startups becoming like the Moody’s or the S and P, the folks who who evaluate the quality of AI for specific use cases because nobody really knows.
It’s really the wild west out there. And one of the big problems, for example, is data contamination where a bunch of these language models are trained on the things that they are being evaluated on. So GPT-four looks like it’s an amazing coder, but it might also just be trained on the data that it’s evaluated on, which means that it’s not actually that great of a coder.
How do we know a yardstick for progress or measurement? What is the right way to approach AI measurement and effectiveness?
So there’s a Stanford Helm project, the holistic evaluation of language models, and these benchmarks where you look on static test sets, how good language models are, how good they are on this static test set. But I’ve been arguing for a long time that that’s just completely wrong anyway, and we need to do something that’s much more dynamic. So ideally, what you want is to see how easy is it for an adversarial person to mess with your model. The harder that is, the better your model is.
And so it used to be very easy to come up with adversarial attacks where these models would just completely mess up. And it’s getting harder and harder over time. So success rate of an adversarial attacker, that’s something that can keep evaluating over time. So we need humans to evaluate these models by trying to break them.
When you say about kind of the adversarial entrance and abilities, does this mean we’ll have, like, an entirely next generation wave of cybersecurity companies around model protection?
Oh, absolutely. Yeah. That’s completely going to change everything. So these models can also get contaminated with data itself. Right? So we need the security layer on the generations of the models. You can do all kinds of prompt injection attacks. These models, right now, they’re still mostly just producing language. Right? But they’re starting to produce code and instructions and actions. And when that happens, then you can mess with the model to get it to produce actions that you really don’t want it to produce, like removing your entire database and things like that.
So that definitely is something that you wanna think about.
My question there is, okay, but is that built by Contextual? Is that built by OpenAI in terms of a data contamination checker, in terms of a health checker? Or is that a next generation Symantec, you name it, security company external that you provision internally.
There are lots of startups looking at this opportunity. Standard security companies are also looking at this right now. So it’s very obviously a very big new threat surface. But I don’t think that the actual foundation model builders, like OpenAI and Contextual, are going to build that technology in house. It’s probably gonna be an external audit.
How much of a concern is data contamination at this stage, you think? If you’re sitting in OpenAI today, how do you think they discuss data contamination?
I think they’re aware of it. They’re starting to invest more in evaluation. They probably should have done that sooner. They have an army of annotators now, human annotators who are checking their models, so they probably have a pretty good sense of how good their model actually is. But obviously, they’re not gonna share that with the world.
I had Yann LeCun on the show who we discussed earlier. Obviously, a very big proponent of kind of open models. Where do you sit in terms of the model that rules for the next five to ten years? And is it different for the model that rules for the next five years versus that that rules for the next ten?
So the way I think about the language model space is kind of as a pyramid. So at the top of the pyramid, we have these frontier models. So these are GPT-four and Anthropic models and things like that that are much better than everything else, but also much more expensive and much bigger than everything else. And then at the bottom of the pyramid, you have open source models. Anybody can train on them. Anybody can fine tune them on their data. That’s a very fruitful area for research.
But I think the most interesting part is kind of the middle piece of that pyramid where you have the most bang for your buck. So that’s from a business perspective, most interesting part where you have mid sized models that have capabilities that you don’t really see at this bottom of the pyramid that you can monetize it in various ways. It’s not gonna be the case that there’s just one model that wins everything. It’s going to be lots of models at different layers of this pyramid being used for different kinds of applications.
So if you have very strong AGI requirements, you probably want to have a frontier model. If you care about it a bit less, maybe you want to have artificial specialized intelligence. If you care about it even less, then you can just take an off the shelf open source model. So there is always going to be a place for open source, but why I don’t think that open source models will move up that pyramid to the frontier is because they’re just too expensive. And this whole flourishing that you see right now of open source models that basically comes from Meta’s generosity in giving Llama away for free.
And if they had done that, then you wouldn’t see that.
We saw Elon’s petition. E mad at stability very much was in favor of it on the show. Where do you sit in terms of Elon’s petition, how did you read it?
This is the petition where he asks everybody to stop working on AI so that he can catch up. A a lot of the narrative in the media right now is really driven by self interest from a bunch of folks in the field. So the whole kind of existential risk debate, I think it actually comes from a very good place and a lot of people are worried about this. And I think there is a non zero probability of AI extinction risk. So it’s something we need to think about, and a lot of smart people are thinking about this, like Benjo and Hinton and all of these folks.
What
makes you
say there’s a non zero chance, though? So there is just a non zero, but very, very, very small chance there will be some sort of paperclip maximizer scenario. Have you heard of this paperclip maximizer? No. Tell me more. So if you give a very intelligent system an instruction like you need to make as many paper clips as possible, then it’s going to turn everything into paper clips. And it’s going to basically destroy the planet and turn everything into paper clips because that’s its objective function. In the process of maximizing paper clips, it will destroy everything else and turn the whole universe into paper clips.
You can see why I’m saying that these are very, very small probabilities. So and so probably the chance of me getting hit by lightning, like, right now is much higher than that happening. I think one of the issues I have with the whole debate around existential risk is that it’s really a tiny probability, but we’re pretending like it’s a massive issue. And I think there are much bigger risks, right, like nuclear war and pandemics and climate change, and those are things we should be focusing on much more.
So the people who are pushing this narrative are really the people who are benefiting from this being the narrative. So these are the incumbent AI companies who want to have either the market regulated because then they benefit because they can deal with the regulation, but small companies like mine can’t. Or they are the folks who benefit from kind of fear mongering in the broader public where people start having a lot of for AI’s capabilities and want to use it everywhere because AI is so smart that it might even kill us all.
There’s a lot of dubious motives behind the scene there.
You mentioned regulation there. I think a big question for me is, like, I don’t think the chasm has ever been greater between private company knowledge, specifically around AI, and then also the regulators knowledge, is significantly behind. How can effective regulation be set with such a large chasm between private sector knowledge and regulator knowledge?
Yeah. We have to invest a lot in educating regulators. The AI community has been terrible at this, and the broader populace just needs to understand much better what AI is and what it can do and what it can’t do. It’s been slightly self interest driven, I think, in that a lot of folks in AI have just wanted to keep the technology for themselves, and that’s why they haven’t really invested in educating the rest of society. That’s really a a huge issue. A bit of a side point there, but I think the people who tend to write the regulation, they generally don’t really understand technology all that much anyway.
I think the main risk so speaking of Europeans, what Europe is going to try to do is overregulate everything and just completely destroy innovation. I’m very worried about that. I think The US is traditionally much better at not overregulating markets to let innovation thrive. I hope it stays that way. Nobody really benefits from over regulation here except the incumbents, and those are the big tech companies already lobbying for regulation anyway and spreading fears of AI existential risk and things like that.
I’m quite concerned about the EU’s regulatory stance around AI. What they’ve suggested so far pretty much makes it impossible for most startups to use any models that aren’t owned and operated by themselves. How do you think about what happens with the EU regulation?
The European Union just has a tendency of really killing innovation because they think they can lead through regulation, And that’s really the only thing that they’re really good at.
It’s a very dangerous thing when everyone wants to be the leader, and the best way to be the leader is to be the first and the strongest and the hardest. And when you apply that to regulation, it’s like It’s even worse, I think. Yeah. Yeah.
Keep going there. There won’t be anything to regulate.
But I wonder what b two b adoption before we do a quick fire. We had Aman on the show, I guess, as I said, and he said, hey. Businesses aren’t really adopting it yet. It’s gonna be tidal wave of adoption next year. So I guess the first question is, what are the fundamental blockers for businesses adopting AI today?
Amanpreet Singh really great at speaking in quotes, by the way. Like, that’s another great quote. Yeah. I I think the tidal wave is coming. There are just big problems that we have to overcome, and these are the things I just talked about. Right? So hallucination, attribution, compliance, up to date ness, data privacy, latency, and I think the whole field is moving in this direction of just making everything ready for enterprise usage. This is gonna happen.
Do you agree with Amanpreet on the timing next year will be the year when enterprises adopt AI at scale, and it will be a freaking train in his words?
I think it’s already happening. I was at a exec event at Google earlier this week, and there were all of these c level folks from all companies across the world, and they were all talking about how they’re using AI. And everybody’s experimenting with it and it’s starting to make it into production already in various places. I don’t think we have to wait for a year. It’s already happening, but there are just some big hurdles that need to be overcome and they will be overcome very quickly.
Do you think it will be a fast or a gradual adoption? I’m always aware that excitement is very quick, but actually, any new technology cycle, it always takes a little bit longer than one thinks, actually.
Yeah. It will be gradual, think. And a lot of work is right now actually going into finding the right use cases for this technology because people are now starting to think that they can use GPT-four for anything. That’s just not true. People are trying to experiment with the right use cases for different types of models.
I spoke to a big, big European company the other day, and they said the biggest challenge for us is we have millions and millions of lines of kind of transaction data. There is no freaking way we are letting any of our transaction data go anywhere off prem. Like, it has to be so secure on prem. It is our lifeblood. How do you think about security of large enterprise data and willingness for it to go into models like OpenAI without security loss?
Yeah. Absolutely. That’s really one of the big questions. Right? And and that’s why what we’re building has this very clean separation between the data plane and the model plane, because then you have much more control over where the data goes.
Can you just talk us through that? What does that separation mean?
Yeah. So traditionally in cloud deployments, you talk about the data plane and the control plane. And so control really is okay. You’re a startup or a company and you want to be able to deploy models to your customer’s VPC, virtual private cloud, but you don’t want any of their data leaving their VPC because it’s their data. So that’s the separation between the data plane and the control plane. So in this new setting where we have language models, there’s this model plane, and it’s kind of unclear where to put it.
So you could put the model inside the customer’s VPC and then you get full data privacy basically, but you have no control over what that model is doing at all. You get no feedback. You get no learning. So you want to find interesting hybrids where you can respect data privacy, keep the data plane inside the customer’s VPC, but put the model somewhere else. So you can do that if you have a decoupling between the retrieval part and the generative part, which is what we are building.
When you look at all the VC fundraisers, do you look at it and go, this is getting crazy?
Not so much. I I think some of the rounds were pretty big, but I think it’s also justified just because this stuff is really going to change the world. And so one that, if it’s right, has massive payoff. There’s this narrative I think in the VC community that there are these crazy rounds happening, but I think they’re happening for good reason. So I haven’t really seen any companies come by where I was like, wow, like, why are they getting this much money? There there’s a few of them where I thought they really have to live up to massive expectations now, they need to actually start making real revenue now.
At some point, there’s going to be a disillusionment with the technology, and then funding might dry up, and then these places are really in trouble.
Can I ask you? When you look at the incumbent sets, who do you think has the strongest strategy execution to date? Is it Facebook? Is it Google? Is it Microsoft? Is it Apple?
I’ve been very impressed actually by how Microsoft has managed to turn everything around by strategically collaborating with a a better AI lab in the shape of OpenAI. And they’ve just really turned that into this narrative where Microsoft is an AI leader, and they really weren’t an AI leader even a few years ago. So that’s been impressive.
On the flip side, who do you think’s not done well and not adjusted to the new landscape?
I’m still very curious to see what Apple will do. They had this great vision of having, like, Siri on your phone and things like that. So that’s, like, one of the first personal assistants. So if that could be a super powerful language model, then that could do very interesting things. But so far, I haven’t really seen interesting things coming out of Apple.
I wanna do a quick fire round. I’ve peppered you with questions anyway. But this is like a more structured peppering with sixty seconds per one. And so it’s much more, you know, informed, and the questions will actually be on schedule unlike the rest of the, you know, the last forty five minutes. Yeah. Sounds great. So what do others not
know that you know to be true? So I I think others are underestimating how early it still is. AI feels like we’ve made so much progress that it’s very hard to enter the market right now, and I I think that that is just not true. It’s still very early innings. We haven’t settled on a lot of things that need to be solved before we can really have this technology be ready. It’s still very early. Me realizing that gives me a competitive advantage, hopefully.
What do you advise founders who are building AI companies not in the Valley? Do they need to be in the Valley?
No. Absolutely not. The Valley is kind of a dangerous bubble in a way where there’s this giant echo chamber happening. And I I think if you look at a lot of great AI companies, they’re not in the Valley and they don’t have to be here. Maybe they should have an office here because there are great universities to recruit from and things like that, but I really don’t see a reason why you would have to be here. What would you most like to change about the AI community?
Hype. I think there’s way too much hype. It would be good if the community at least acknowledges that and tries to really pay attention to the things that matter, like having technology that actually works and not just jumping on the next hype train and there’s an auto GPT thing that is going to change the world but doesn’t actually work. So there’s there’s a lot of debate right now that is just driven by Twitter and just sound bites and quotes like like Amanpreet, where I think it would be better to be a bit deeper and think a bit more carefully about what we’re doing.
Some fantastic quotes though, weren’t they? I mean,
credit They were. I’m with you totally. Okay. Do you agree that some of the biggest businesses to be built in AI over the next years will be services businesses for large enterprises helping with AI implementation.
Yeah. Totally. I don’t know if those are gonna be new businesses or existing incumbents. So the hyperscalers are also trying to play that role. There’s just so much demand right now for AI in any kind of enterprise, and it it’s still very hard to get it right. So a lot of these companies are just looking for help, and so there are just opportunities there. How do AI and
philosophy
help each other in your day to day role? I’m very happy that I studied philosophy. So philosophy is really about conceptualizing anything and any arbitrary level of abstraction and that ability you can use anywhere. So for AI in particular, I think that philosophy and AI are an interesting combination because philosophy is about the stuff that you can’t really do science about yet. So at some point, the things that people are philosophizing about now, they become scientific questions that just have answers or hypotheses, then they’re no longer philosophy.
Right? So natural philosophy that used to be a thing, we now call that physics and mathematics. And I I think with AI, there are lots of questions that we still don’t even really know how to ask yet, and philosophy is great for thinking about those kinds of questions.
What’s the strongest
belief that you had which turned out to be wrong? The strongest belief is that I really underestimated how important scale is in artificial intelligence. And I think this is really one of the things that OpenAI has excelled at. If you throw an order of magnitude more compute and data at AI systems, then they just become much, much better. And if you keep that scaling up, you have these scaling laws that we know about now. I really underestimated this. And for a long time, was just saying like, oh, yeah.
Look at these silly OpenAI researchers. They’re just scaling things. They’re not inventing new algorithms. That’s not cool. And I was very, very wrong.
I’m gonna apologize in advance for this one. Okay? I’m sorry. I’m asking it anyway. What do you think the timeline is for superintelligence?
So so despite Nick Bostrom’s book, I still think that superintelligence is actually very ill defined. In many ways, we have already achieved superintelligence. And so in the fifties, we achieved mathematical super intelligence. So computers in the fifties were already better at calculating stuff than humans. I don’t think that that really is a well formed question. And if you’re asking about AGI, and I think AGI itself, a lot of people this is a mistake I made where I thought AGI kind of meant artificial consciousness or something like that, which also doesn’t really have a meaning.
But if you look at how OpenAI and Anthropic and these places define AGI, systems achieving capabilities that allow them to effectively do the work of humans for the majority of economically valuable human tasks, then we’re not that far away. And so I think in the next, like, five to ten year, that sort of economic displacement is likely to happen.
What’s the most painful lesson that you’ve learned that you’re also pleased to have gone through?
I think when I was young, I maybe put ambition before people sometimes. I learned this the the hard way, I think, where I just didn’t have enough empathy for the people I worked with. And as I grew older and more mature, I I realized more and more just that it’s really all about people and working with fantastic people and and doing cool things together and and changing the world together. So I’m very happy to have learned that lesson because that makes me much better at my job right now.
Final one. Ten years time. If all the stars align, where’s Contextual then?
So if all the stars align, OpenAI and Anthropic and all of these places, they had this great first generation technology. So they’re kind of like the LeCos and AltaVista of search engines. And the technology we have is more like PageRank, and that would make us the the Google of language models with the right technology at the right time and with the right execution. So that’s what I would hope for.
Dou, I’ve absolutely loved doing this. I can’t thank you enough for putting up with my wayward, sometimes very naive questions. I can’t thank you enough for letting me invest in you, and I really appreciate the time today, my friend. Thanks for having me. What a show. I have to say, I just love diving into AI with some of the best in the world. But if you wanna see more from us, of course, you can on YouTube by searching for two zero VC. But before we leave you today,
· Sponsor read0 min · 540 words
did you know that over 50% of your day is filled with tedious tasks? What would you do if you got half a day back? Well, you can with Coda, the all in one platform that changes the way your team works together. And Coda just introduced an AI powered work assistant to take the busy out of work. With Coda, your team’s important workflows and content already live in one place, and Coda AI helps each team member focus on the highest priority work even as priorities shift across the team.
By taking over recurring, and let’s be honest, very tedious work often, Coda AI empowers team members to prioritize longer term and strategic work to accomplish goals. Coda AI not only makes space for collaboration in each person’s workday, but can make it easier to stay in the loop and share informed opinions. So if you want to work a system that lets you get back to work, you can get started with Coda AI today for free. Head over to coda.io/20vc, that’s coda.io and get started for free. Coda.io/20vc.
And speaking of tools we cannot live without like Coda, we have to talk about Brex, the all in one financial stack trusted by founders. Founders have to think globally in order to open new markets, unlock cost savings, and gain access to new talent. That’s why having the right financial stack is more important than ever, and that’s where Brex comes in. With Brex, you get a high limit corporate card, a high yield business account with up to $6,000,000 in FDIC protection and bill pay, all build with a global first mindset.
Brex enables you to operate in more countries and currencies than any other provider, so you can pay vendors, run payroll, and make international payments faster. We all know that’s a must have for startups at any stage of growth. So are you ready to learn more about Brexit’s global first solution? Visit brex.com/20vc. That’s brex.com/20vc. And finally, AngelList is fast becoming the center of the venture ecosystem. The startup world is just buzzing about how fast they’ve been shipping products that meaningfully improve the lives of both startups and fund managers and their investors.
For startups, AngelList reduce the friction of cap table management, banking, and fundraising all in one place. Teams can focus on scaling and let AngelList handle the rest. Thousands of start ups have moved their cap tables to AngelList in the past year. AngelList also supports large venture funds and their teams with an automated software first approach and the best customer service in the industry. Fund managers can focus on making great deals while AngelList handles reporting, taxes, compliance, and more. So if you’re ready to scale your startup or fund with the platform at the center of it all, visit angellist.com/20vc to get started.
Now next week, we are taking a week off. Yes. I am going away with my family to the British Coast, so we will not have any shows next week. And this will be a first in a very long time, but even I need a rest sometime. So I so appreciate your support and stay tuned for more episodes in ten days.