Cold open
I don’t think Sam Altman has done a service to the world by talking about how close AGI is. I think he has made several predictions now that are wrong and that were obviously wrong at the time he made them. AI is gonna kill the whole world in two years. He did a world tour where he spoke to every major leader the world over to tell them, hey, this technology is gonna pose an existential threat. And I think that was academically disingenuous, and I think did a disservice to the technology he loves. A lot of the world does not think scaling laws are super prevalent.
This is 20 VC
Intro
with me, Harry Stebbings. And today, we are joined by Nick Frosst, Canadian AI researcher and entrepreneur, best known as the cofounder of Cohere, the enterprise focused LLM who has raised over $900,000,000, most recently raising a $500,000,000 round, bringing their valuation to 6,800,000,000. Today, we discuss how on earth they compete when competing against the billions of dollars that OpenAI and Anthropic have. Cohere has hit a 100,000,000 in enterprise ARR. And before founding Cohere, Nick was a researcher at Google Brain alongside the incredible Geoff Hinton. But before we dive into the show today,
· Sponsor read0 min · 605 words
I love seeing the team come together to make this show happen. What I don’t love is trying to keep track of all the information, the data, and the projects that we’re working on across dozens of platforms, products, and tools. That’s why we use Coda, the all in one collaborative workspace that’s helped 50,000 teams all over the world get on the same page. Offering the flexibility of docs with the structure of spreadsheets, Coda facilitates deeper teamwork and quicker creativity. And their turnkey AI solution, the intelligence of Coda Brain, is a game changer.
Powered by Grammarly, Coda is entering a new phase of innovation and expansion aiming to redefine productivity for the AI era. Whether you’re a startup looking to organize the chaos while staying nimble or an enterprise organization looking for better alignment, Coda matches your working style. Its seamless workspace connects to hundreds of your favorite tools, including Salesforce, Jira, Asana, and Figma, helping your teams transform their rituals and do more faster. Head over to coda.io/20vc right now and get six months off the team plan for startups for free.
That’s coda, coda.io/20vc, and get six months off the team plan for free, coda.io/20vc. And while coda keeps the engine running smoothly, let’s talk about Brex, the ultimate financial stack for startups. So when Brex was founded, it wasn’t just about creating another financial product. It was about solving the really gritty challenges that founders face daily. Let’s be honest, building something from the ground up is hard enough without dealing with clunky outdated banks that pile on fees and leave your cash idle. Brex is different. It’s the financial stack that scales with you no matter where you are in your journey.
From corporate cards to maximizing your runway to earning yield on your cash. Brex was designed with founders in mind to make every dollar go further so you can focus on building. And here’s what really stands out to me. Brex combines the best of checking, treasury, and FDIC insurance in one powerhouse account. You can send and receive money globally at lightning speed, earn yield from day one, and still access your funds whenever you need. Plus, with 20 x the standard protection through program banks, your cash is not just working harder, it’s working safer too.
It’s no surprise that one in three venture backed startups in The US with companies like Anthropic, Coinbase, and Robinhood. I mean, my god, these companies are incredible. Trust Brex to help them grow. If you wanna join the smartest startups on the planet, head over to brex.com/startups and see what they can do for you. And speaking of incredible companies, don’t forget what really keeps those customers coming back. Trust is the ultimate currency in business, and today customers expect it faster than ever. And that why over 10,000 global companies trust Vanta.
Vanta automates up to 90% of the work for in demand compliance standards like SOC two, ISO 27,001, and more using smart AI to centralize workflows, manage risk and get you audit ready in weeks, not months. So you can stop chasing paperwork and start closing deals. And a new IDC report found that Vanta customers achieve 5 and $30,000 per year in benefits. That’s insane. And the platform pays for itself in three months. I had no idea about these. Whether you’re growing fast or just getting started, Vanta connects you with trusted auditors and experts, support to help you build trust with customers.
Get a thousand dollars off your first year at vanta.com/20vc. That’s vanta.com/20vc. You have now arrived at your destination.
Conversation
Nick, I’m so excited for this dude. When I had Aidan on the show, he was like, you’ve gotta have Nick on. He’s the real star of the show, and he introduced us way back then. Mhmm. So I’m so excited that we can make this happen. Yeah, man. I’m happy to be here. Now before we dive into Cohere, I have to ask, you were Geoff Hinton’s first hire at Google Brain. Mhmm. And so then you’re put in a room with Geoff Hinton, you get to work with him every day.
What was the biggest lesson from working with Geoff, a legend of the industry?
Yeah. I I learned yeah. I I love working with Geoff. I learned everything I know about research from those those I think we were there for four years, three years. I was very surprised at how creatively and playfully he approaches research. When we would discuss, like, algorithms or or or, like, optimizers or or loss functions, we would discuss them through physical analogy. So we’d spend a lot of time talking about, like, imagine there’s, a ball here and an elastic band to this thing and a pulley here and this is what the, you know, it’s on this kind of a surface and a lot of it was descriptions in the natural physical world and that was very, yeah, like playful.
And a lot of it was approached with like, oh, what would happen if, you know, with curiosity? And I didn’t expect when working with him, I didn’t expect that. I expected it it to be much more like, you know, just here’s the equation. Let’s let’s figure out what the derivative is and let’s let’s go from there. Whereas instead, lot of it’s based on, like, intuition.
When you look at Google Brain and you look at DeepMind, a lot of things that really kind of Google were asleep at the wheel given them not being at the forefront in what was the consumerization of it with ChatGPT. Mhmm. Do you think that’s fair?
I don’t know. I mean, it’s certainly interesting. Look, like, the transformer was invented at Google. Right? Like, there was in 2017, Aidan, amongst with many other brilliant people in Google Brain, published the transformer as an architecture. It wasn’t then commercialized very quickly within Google. It wasn’t scaled up very quickly within Google. A lot of that work had to be done elsewhere and years later. So that’s interesting. Why that is? Like, what what systems are in place to make that be the case? I I don’t know.
I will say there’s still a ton of brilliant people in DeepMind, I think, now. It’s just consumed the rest of it, doing great work, and they continue to make good products. It is interesting that all the people who worked on the transformer left to continue to work on the transformer.
For people who don’t know, and just to set the scene before we dive in properly, what is Cohere and how does it differentiate from more generalized models that are maybe more well known like your OpenAI and your Anthropics?
Yeah. So we’re we’re a foundational model company like those other two. So we we build foundational models. We build language models. I don’t know. Maybe there’s, like, 10 companies in the world that are building large language models. In the West? There’s a few that have popped up recently. Some number less than 20 in the whole world. Most of them in America, a handful of them in China, Us in Canada, and one in France. So those are really the the companies out there. We’re unique singular focus on bringing this technology to enterprise.
We train a model that is good at enterprise tool use. So, like, we train a model that you can, you know, give it a bunch of tools and APIs within your business, give it access to your business’s data, and then you can ask it to help you with something in your work and it does a good job of it. So that’s what we train
it for. How does a focus on enterprise over consumer change the way in which you train and build a model?
Yeah. So the models themselves, like transformer architecture, which is the original model that, yeah, that was introduced in 2017, hasn’t changed very much. The whole industry is still using transformers. We’ve changed the way we train them, but the model architecture itself, you know, we’re approaching ten years of the same model architecture. When we train our model, we’re not training it to be like an amazing conversationalist with you. We’re not training it to, like, keep you interested and keep you engaged and occupied. We don’t have, like, engagement metrics or things like that.
We’re just training it to augment you in the workplace. We’re just training it to help you do your job. And that means the type of data we train it on is very different. So recently we started doing a bunch on, like, synthetic data, generate a whole bunch of data to create, like, fake companies and fake emails between people at these fake companies and fake APIs within those fake companies. And then we train the model in that synthetic environment to help out within that fake business.
Do you think data is a bottleneck given the ability for synthetic data to produce infinite supply?
Yeah. Data is still a bottleneck. You need real world data in order to start a process of synthetic data. Synthetic data has helped a lot, and it’s made models better than they would be if they didn’t have getting access to high quality data is still something people think about. We still, you know, we still make a whole bunch of data in house with annotators who are making real data and not synthetic data.
When you think about kind of the three pillars of compute algorithms and data Mhmm. Which one do you think is most constrained or the biggest bottleneck? It’s interesting. I mean, the algorithms haven’t
changed very much. They’ve changed a little bit. You know, when we started this industry, originally, we were just training base models, which are called not base models at the time. They were just called large language models, but they weren’t trained from human feedback. So all they would do is, you know, take in the first part of a sentence and write the second part of a sentence. But if you tried to have a conversation with them, it, like, it wouldn’t work because that wasn’t the data they were trained on.
Since then, now we we train models in a few different steps. There’s like a base modeling step, then there’s a reinforcement learning step from human feedback with SFT data. After that, you know, there there might there’s a a variety of other reinforcement learning techniques you can do. But the algorithms, I I think, are not the bottleneck in terms of making those models more useful. A lot of it is still getting good quality data and then making good quality synthetic data from your good quality real data.
When we think about kind of the bottlenecks, bottlenecks, that leads to potentially a plateauing that people are worried about. And everyone seems to now be on the train of, hey, more compute, scaling laws are more real than ever, and we will continue this exponential progress with more compute. Do you agree that we are seeding the benefits of scaling laws for the continuous next twelve to twenty four months? Or do you think that actually more compute will not just lead to more progress?
Well, how much better do you think GPT-five was than GPT-four? I actually think it was worse. Tells you something about the nature of just throwing more compute at the problem.
Does it or does that show an so why do I think it was worse? I think it was worse because actually the way that they now do my model selection is slower and more cumbersome and actually it’s pain. It gets it wrong sometimes to me. I just want a quick answer and it suddenly goes into deep research. I’m like, for fuck’s sake, just want a quick answer. I’m like, yeah. Alright. PhD, calm down. Like, do you know what I mean? And so, yeah, I think it’s a worse product in that respect.
And I think we waited for a year or a year and a half for model auto selection.
I I think, like, if I go back to your your your original question of, like, you know, do I do I think just throwing more compute? Like, some people are thinking there’s a plateau. Do I think there’s more compute? Like, I think we need to agree on where we think the technology is going to establish whether or not there’s a plateau. Language models are incredible. I use them in my work life as often as possible. One of the reasons why we’re focused on the enterprise is because that’s really where I think large language models are useful.
Like, if I look at my personal life, there’s not a ton that I wanna automate. You know, like, I actually don’t wanna respond to text messages from my mom faster. I wanna do it more often, but, like, I I wanna be writing those. I wanna be, like, engaged, you know. Whereas in my work life, there’s a ton of stuff I don’t wanna do. Like, we need to get to a stage where I can, you know, open up north and I can say, hey, file my expenses.
And then it can figure out, okay, cool. I gotta, you know, look through all your emails. I gotta look through photos of receipts you’ve taken. I gotta cross reference that with the things you’re allowed to expense via internal documentation. Then I gotta figure out what the API is for how to expense things within your company, and then I gotta do all of those and get approval before I do them. Like, that’s a super that’s a many step process, but that’s where the technology is going. That work, the work of making a model do that is not plateauing.
That’s more modeling work. That’s more product work. That’s like building better connectors. That’s building safer data integrations so that you can trust giving a model access to the types of stuff I just said. That stuff’s still ongoing, and that’s what we’re we’re working on. I think when people are talking about building towards AGI, like, don’t think this technology gets us there. When you say gets us there Yeah. What is there? Well, yeah, great question. We’ve had many years of people discussing AGI and not many definitions thereof.
Like, next to none. I mean, my
definition is when Sam Altman and Microsoft decide.
Yeah. They’ve changed their definition a few times on that. When I say AGI, what I mean is a computer that you treat like a person. When you use a computer and you expect it to behave like a person and treat it that way, I’ll call that AGI. Do you not think we’re already there then? People do not treat language models like they treat people.
Do you think OpenAI and Sam Altman then now realize that more compute does not lead to this exponential progress when they look at GPT-five?
I don’t know. Yeah. I don’t know. I think, like, that, you know, they’re a great company. They build a really cool consumer product.
Why does the world still think scaling laws are so prevalent when you don’t? A lot of the world does not think scaling
laws are super prevalent. If you go out into a university and, like, talk to the students there who are studying computer science or even the students who are not, and you ask them like, hey, you know, is throwing more compute at this problem gonna get us to AGI? Most of them say no.
You you mentioned there a couple of different use cases in terms of, like, expense management was one that you clearly articulated. A question that I think I have and a lot of people have is how far do models go in terms of value capture and see application layer? And you’re seeing Anthropic now with Claude, really challenge the Cursor of the world, you’re seeing OpenAI with a lot of consumer products, you with a lot of enterprise use cases. How do you think about whether they stay as AWS style commodity layers or whether they extend into value capture of application layer?
Yeah, that’s a good question. I don’t see the two as that different. I see the two as related. And if you wanna be making a good product with a large language model, you are best suited training that large language model for that product. That’s one of the really interesting things about LLMs is that they’re really phenomenal. They generalize really well, but they don’t generalize as well as you might think. And if you wanna make the best model for a given interface, it’s best to be training the model on that interface.
So I think the two things are more related.
So do we see this deeply specialized unbundled model world where you have exactly that? Very, very specific use cases where models are trained for and 11 labs of the world would be another brilliant use case with, you know, specifically voice. Is that the world that we live in?
There’s like a spectrum. Right? Like, the old world of machine learning back in, like, the old world back in, like, 2015 or something like that. When the world was new. Any task you wanted to do with a neural net was the best neural net you were gonna get was training a model on that task. So if you wanted to make a task that was gonna, like, a neural net that was gonna identify pictures of cats and tell you how many cats were in an image, you were best suited to train a model on identifying pictures of cats.
That was, like, the best. And the the world of machine learning and the first half of that decade was all about, here’s a problem. Make a dataset, train a new model from scratch, or maybe take, like, sift features or maybe take, like, whatever, some base model, but pretty much fine tune like, train a model on the dataset itself and go to production with that model. That’s not the case with language. If you wanna make a model that’s the best at, like, summarization, you can’t just train it on summarization.
You have to train it on all language. That is the technological reality that has brought us to where we are today. And that’s true for, like, foundational models and not true for the neural nets of of 2015 and and before. That’s super interesting. But that’s like a spectrum. Right? Like, on one end is every single task, train a single model for it. On the other end is train a model, one model to do everything. I think the reality of what we’re seeing with transformers is they’re not at this end of the spectrum.
They’re like a little over here. And they’re like, train a model that is generally good at all language and refine it on the type of stuff you wanna do with it. So you’re seeing that like Anthropic’s code models, very good at code. But they didn’t train a model specifically like a refactoring model or a debugging model or they didn’t train a model just writing test cases. Right? It’s a model that is good generically at code. For us, for Cohere, with our focus on enterprise and, like, secure deployments and customizations for our and for our customers, that means training a model that is good at helping people in an enterprise setting.
So, like, using internal tools, reading through massive amounts of documentation,
understanding Does that mean the model doesn’t need to be as good? Again, I’ve learned to be incredibly blunt. If someone wants to criticize, they say, oh, well, if you look at models or evals, Cohere is not as good. Well, there’s a whole long conversation to be
had about evals. But effectively, like, what we care about is not discord, not the hype, not the discourse. Like what we care about is if a customer uses our model and they try to do something with it, we care that it works as easy as possible. Like, that’s what we optimize for. None of those are really reflected in the various benchmarks that cycle through every year. And so we don’t, like,
focus too much on that stuff. Do you think the benchmarks are bullshit? Because we place a lot of emphasis on them in on Twitter sphere, on the Reddit sphere. Are they bullshit or are they accurate reflection of model progress?
Let’s go back in time a little bit. When we first started in this industry, the benchmark that was used the most was called LM one b. That was a benchmark that was like taking in the first part of a text, like of a of a newspaper and then writing the second part of a of the newspaper article. After that, there was a benchmark called Hellaswag. Do you remember that one? Yeah. I do remember that one. Alright. Cool. So that’s like 2022. So that’s like That’s my intro.
Yeah. Alright.
Cool.
No one’s talking about that anymore. Right? Right. A lot of people talk about like AIM as like a math reasoning, AIMY or actually don’t have math reasoning benchmark. None of our customers ask the model to do math reasoning. That doesn’t come up in the workplace that often. That comes up in a few workplaces where mathematicians work, but there aren’t a ton of people out there making a living doing math reasoning. Stuff like the ARC AGI challenge is a benchmark that people talk about, but that’s like a pixel manipulation challenge.
It’s like, you know, taking in like a grid of pixels pixels and based on rules predicting the next one, that’s not a thing any of our customers have ever asked the model to do. So do I think they’re all bullshit? It’s interesting. I don’t know. There’s there’s good scientific work in some of them. I think it’s very interesting to evaluate emergent capabilities from models. But they’re not
an actual reflection of the utility value of models.
They’re a reflection of how much the model have been trained on those benchmarks. So you can gamify them essentially? Oh, you can definitely gamify them.
Yeah. Do
the
big players gamify them?
I don’t think those leaderboards are that helpful. I think in a consumer space, it’s cool. I think if you’re making a consumer app and it’s, like, exciting and fun and people like to look at it and they wanna try out the most recent thing, that’s fun. That’s cool.
Given the pace of deployment, we are seeing model evolution so fast and so rapidly that you’re essentially seeing this kind of decay rate on models being greater than ever because it’s like next one, next one, next one. And actually, they’re still being trained though on h one hundreds or NVIDIA chips from eighteen months ago. Is there a misalignment in terms of the progression of models versus the progression of chips? You can cycle through
new versions of models quicker. I mean, it’s still it’s very slow. Still, like, when I was training neural nets in 2000 and, I don’t know, ’11, and it would take like hours to days. I remember being like, this is crazy. I can’t believe this takes so long to train this model. Now we spend months, months training models. So, like, you know, that’s that’s a timescale I I didn’t anticipate when I was working on this a long time when I was working on neural nets a long time ago.
But that’s still very different than the timescale of working on chips. Right? Like, that’s still slow. I think when you talk about, like, we’re seeing all these models iterate so quickly. Like, yes, on the one hand, we’re seeing models iterate really quickly and people are releasing new models. On the other hand, there’s still the transformer that was invented in 2017. And there’s still sequence models, and they still take in words and predict the next word. And we’ve changed how they’re trained a bit. We’ve added on steps, like now there’s modeling step, then a SFT, like supervised fine tuning from human feedback where, like, somebody writes a sentence and then writes the second the response they want and we train on that.
And then there’s a reinforcement learning aspect where the model is generating and you’re telling it that’s good, that’s bad or something. So there’s, like, new ways of training it. But fundamentally, the tech is still the same. We keep making them better, keep iterating on them. But it’s not as though we’ve, like, anybody has, you know, trains a model that’s fundamentally different than a transformer. It’s an interesting dichotomy. Like, on the one hand, there’s constantly new stuff. On the other hand, we’ve been working on the same stuff for a while.
We have been working on the same stuff for a while. The thing that has seemingly changed is the value of the people working on the stuff. You know, we’re now seeing billion dollar people in terms of Zuck’s willingness to pay for, like, chief scientists. Joel Pineau. Yeah. Joel. Yeah. Joel from Facebook. Yeah. Yeah. Or Meta. My question to you is, how do you think about the war for talent that we’re seeing today?
I think there’s a lot of crazy headlines out there.
I don’t I don’t know how much of it is real. I know You don’t think it’s real that Anthropic are paying $10.15, $20,000,000 for great AI researchers? I I have no idea.
I know that there are lots of people who are adding that much value and there’s lots of people who are, like, bringing that much value into the industry. It’s a super impactful industry. So I know that there are people that are bringing that much value. And I know that there’s lots of brilliant people and I know that it’s really demanding work. It’s really hard work. It requires a lot of experience, a lot of ingenuity, and a lot of dedication. It’s a good place for for people to be spending their time.
And I’m think it makes sense that many of them are rewarded very well. That being said, like, when I see the stories of, you know, Meta hiring people for, like, a 100,000,000, like, I read as many stories of those as I read of people leaving the next day. So I don’t I don’t know what’s going on over there. But what is that 5,000,000 on an AI researcher? There are certainly lots of people who, through our equity, own what you’re talking about.
Do you worry that the industry is becoming commoditized or transactionalized with the hype around it? I don’t like Yeah. I do think
the hype around it is misleading sometimes. Like the technology I’m in such a strange place of being caught between the technology is the most beautiful transformative technology I’ve ever worked with. It is already fundamentally changing the way I do work, and I’m very sure it will fundamentally change the way we all do work soon. On the other hand, there’s a lot of hype around it. There’s a lot of misleading rhetoric. There’s a lot of misinformation. And I don’t think the hype is necessarily helpful for getting to the truth.
Can I just dig in there? Yeah. What do you think is the hype and misleading rhetoric that is most damaging or confusing?
Yeah. I think the hype around AGI is the most damaging and confusing.
This assumption that we will all have no work to do, we’ll all be in UBI. Yeah. This
isn’t really in the discourse as much this year as it was last year and the year before that, and that’s because it’s pretty clearly not true. But the idea that, oh, this technology is, like, poses an imminent existential threat to humanity was incorrect and not helpful for talking about the real ways in which this technology could be damaging, the real ways in which this technology could shake up a system and cause rapid changes. It was not helpful for getting people to understand what the technology is.
Right? So I I don’t know, like, don’t hear that as much anymore these days. I think that’s because people have realized that that’s not the case. But the remnants of that discourse are still in the world, are still out there. Yeah.
I think the remnants are, and I think they’re most prevalent in the way the internal employees in large organizations respond to AI being introduced. People do not welcome the introduction of AI in large companies, maybe a European son, but a lot of very nervous and scared and do not embrace it wholeheartedly.
Yeah. I think we we I haven’t found that as much. When we’ve worked with our customers, I find a lot of people are are interested and excited about using an LLM. Mostly that’s because I think they realize that the LLM is augmentative for the most part, and it it allows them to not do the things they don’t wanna do. The artistic
I mean, it’s in a nice way. Do you actually buy that? Do I actually buy Yeah. Like, nice. Yeah. Yeah. I mean, I have Benioff on the show from Salesforce two days ago and he says, like, oh, the same human plus agent. Yeah. Yeah. Are you serious? Most 25, 26 year old marketing managers or SDRs, I’m sorry to say it, they’re not brilliant. They do not love the craft. They are not better than a phenomenal agent will be in the next twelve months. They will be replaced.
Oh, no. Yeah. I I actually believe that. Yeah. No. I believe what I said. Sorry. Yeah. Yeah. I fundamentally believe that this technology that there are things that it’s way better at you, better than you at, but there are still lots of things that people are better at. Like, l look, LLMs are incredible. People people have been using them, you know, for years now. There has been no independent breakthrough that an LLM has made. Nobody has seen hey, nobody asked an LLM solve this problem no one’s solved before and get the the answer.
The breakthroughs are still people. Is that
not only a matter of time?
No. That’s not a matter of time. That’s that’s fundamentally the the way that sequence models work. Like, when you’re training statistical models of text, they’re phenomenal. They are capable of generalizing across unseen tasks, which is why they’re so useful. But when you’re talking about, like, the 25 year old marketer or something, some portion of their work is like writing text on a computer. It’s like all the information’s out there. They have this document and that document, this tool, that API. They just need to take that and turn it into another form and, like, you know, combine it and then put it out there.
Like, that’s the work. That’s some portion of it. That’s not the majority of it. Most of it is talking to people, like understanding the culture, understanding the zeitgeist, understanding, like, what’s gonna hit, what’s relevant, using their intuition and their human experience to understand how they can be helpful or what they can do. And that is not in the dataset of text from the Internet.
I disagree because you know what they do now? They go, hey, I’m running a campaign for Avian, come up with three different story lines that would be cool for us to do. Make it relevant to news cycles today Yeah. Is the prompt. And then they come up with three, and they’re like, oh, that one’s pretty good.
And I think that sounds like a good usage of it as a starting point. And then I’m sure it is worked on, and some of those are thrown out because of things they understand. And some of them are like, oh, that’s a great insight. I’m gonna run with that. I’m gonna tinker that. Like, what you just described is a good use case, and I can imagine as a jumping off point, that would be helpful. But that’s not where the work ends. Like, that’s the beginning of the work.
You’ve just you’ve just augmented, you’re now starting with, instead of a blank page, you’re starting with things to go for.
So you don’t mind that we’re gonna see this dramatic reduction in team sizes? I think we’re gonna see changes
to the nature, like to the workforce. In the same way that we saw changes to the workforce when the computer was created, when a personal computer, when the internet was created, when the printing press happened, right? Like we’ve seen, like in the industrial revolution, right? Like we’ve seen drastic changes to the workforce and and that’s gonna keep happening.
What will the changes be? Like, what will a company look like, do you think, in five to ten years? That’s that’s an interesting question.
I think you will arrive to work and you will sit in front of a computer, and you will predominantly use language to interact with that computer. Anytime there’s something to do that you know can be done and you know, like, the information is out there, doesn’t require creativity or insight, you know that it’s there. You just need to do it and doing it’s kinda boring. You will get the model to do it for you. That mostly looks like sitting down, speaking to the computer, getting it to do the things you don’t wanna do, and then spending your time talking to other people, thinking about how it can be useful, whether or not what it did was good.
I think that shift is maybe chaotic. And I would like us to be spending time thinking about how can we make sure that change is as easy as possible? How can we make sure that making language models allows people to do the stuff that they’re good at and that they like? How can we make sure, like, the labor force is resilient? How can we make sure, you know, that income inequality doesn’t go up as a result of that? Like, those are the types of things I would like us to be talking about, and those are the important things.
I want to go back to your earlier question, like, when talking about, you know, the eight risks of AGI, like, I think those existential threat questions made it harder to talk about the real things, you know, like income inequality.
Do you think AI does more to help or to hurt income inequality? I think it depends on policy.
I think if there’s good labor policy, I think it could help. I think if there’s bad labor policy, it could hurt.
Can you explain that to me?
Look, when you saw the last industrial revolution, broadly speaking, everybody looks back on that industrial revolution and says that was a good idea. Like, nobody’s saying, hey, we shouldn’t have automated. Before the revolution. It was, I don’t know, 90 something percent people working in farms. Now it’s, like, five, less than that. Everybody thinks that was a good idea. It was a crazy time. And if you read stories about what went on during that moment, there’s lots of things that people did that they stopped doing pretty quick, like having kids work in coal mines.
That was a crazy thing. And out of that industrial revolution came a whole bunch of really good labor policies, came unions, workers’ rights, things that I think we also think were a good idea and resulted in not only, you know, better lives for people, but in actually more productivity, like actually a better economy, actually, you know, a better world. A lot of those were were from public policy. A lot of those were from things created in unison between, you know, businesses and governments.
Why do we need to have such significant policy change if it only augments humans and it doesn’t replace them?
Right right now, we’ve seen income inequality go up over the past several years. And a lot of that was happening before AI, before language models were popular. And I’m worried that technology has the potential to exacerbate that without being deployed correctly and without having good policy around employment.
Can I ask you, when we think about like problems to solve, I think a lot of people also get worried about the open versus closed argument? How do you feel about where the future of efficient AI lands in the balance between open versus closed models?
So at Cohere, we make our foundational models and then we release the weights for non commercial usage. So we’re somewhere in the in the between the like open and closed. Right? We’re a for profit company. Like we we exist to make money. We release our weights for scientific and research and, like, you know, you can download it on your computer and run it. That’s a good sweet spot for us as a business. That allows us to, like, you know, build credibility within the community. If people wanna check out our weights, like, can go check them out.
Right? Like, there’s lots of companies that started out as open who no longer release the weights of their models or who never did. Right? So we have our models out there. You can go look at them. You can use them. You can validate. Hey. Did they work on my problem? Yes or no. But if you’re using them for commercial purposes, you gotta talk to us. And then we figure out a commercial relationship so that we can, you know, exist as as a business. That works for us.
I’m surprised there aren’t more businesses taking that tact. Yeah. And more foundational models taking that approach.
Do you think Meta will move to a closed model approach from an open?
They’ve certainly hinted at that. Right? It certainly looks like but I don’t know what they’re I don’t know what they’re doing over there. Yeah. I don’t I don’t think a lot of people know what they’re doing over there. And I don’t spend a lot of time thinking about it.
Do you not think it’s helpful for founders to be very aware of competitive landscapes in case they’re ausked about them by customers, in case customers are going, hey, why aren’t you more open? Why aren’t you more closed? Are we is our data secure if you’re As with most things,
like a middle ground is the right place to be. Right? You could spend your whole time as a founder only looking at competitors and being like, oh, why are they doing that? Why are they doing this? What’s going on with that? You know? You know? And that will, I think, not be helpful for you. You could also spend your whole time, you know, with your head in the sand only thinking about what’s going on in your company, and I think that would not be helpful either.
You have to find some middle ground. The discourse around AI is inescapable. You would be hard pressed to ignore it. It is every other headline. I don’t think there are many people who work in the industry who suffer from not enough information about what’s going on in AI. Right? I think there’s a lot of people who suffer from way too much of it and obsessing over the minute details of, like, so and so got point 2% better on this thing or, like, you know, is constant small changes in businesses out there.
And I think that can mislead you from staying grounded. What are you actually doing? Who are you actually helping? How is this making, you know, things better for your customers?
Do you think we will still have prompting as the core user input guidance mechanism in five years time?
Prompting as in like you write something to a model and it writes back? Yeah. Yeah. Yeah. What else would it be?
The way that it changes, the way that you do it changes. You wouldn’t say, hey, make it a funny tone or hey, add in a light personalized style that also is sincere. I think the
idea of prompting as a skill will become less relevant. And if you look at, like, that’s the trajectory. Like, I started doing this, if you wanted to get a model to summarize something, you wrote the first paragraph and then you wrote in summary colon, new line, and then you generate it. That was the skill of prompting, was figuring out how to trick a model into getting it to do what you wanted to do. And that’s because they weren’t trained on feedback from people, they were only trained on text from the web.
And so all they were were sequence models from text on the web and nobody on the rep on the web wrote, please summarize this for me, and then a summary. They wrote a paragraph and then wrote, in summary, colon, and that’s if you wanted to get the model to do that, that’s what you would do. Language models are like, we’re training them more to fit how people expect them to work. And that means that getting good at prompting is less important. So I think the idea of saying like, oh, yeah, you gotta learn how to prompt is gonna go away.
I think the idea of saying you need to learn how language models work, then you need to know what they can and can’t do. In the same way you had to learn how a computer works and what it can and can’t do. You had to learn how a telephone works and what it can and can’t do. Like, I think that’s gonna exist. And and that means, like, prompting is gonna exist. The idea that, you know, you write to a you write something to a model or you say something to a model and then you get the response back.
And if it’s not what you like, maybe you iterate a little bit like that’s gonna exist. That’s fundamentally how the technology works. But the idea of it being a discipline that you have to train to do, we’ve already seen that trajectory. Yeah. It’s it’s already gotten easier. I look for people who know how a language model works. You know, one of the things that’s been a necessary component of a working Echo here is you can’t think the technology is magic. You can’t think this is like we’re doing spells.
You have to know how a language model works, how it’s trained, and what that means for it. What emergent capabilities happen, which ones don’t. You can’t think, oh, yeah, I just asked the digital god to do my work and then it does. That’s not what the technology is. And thinking that will not help will not help you build it and it will not help you use it.
If we, like, hone in a little bit more and you you led or were a large part of the latest fundraisers we were chatting before. When you think about the fundraising journey, how was that journey? And are there any big lessons from it was 600,000,000 there.
Yeah. I actually quite I quite like talking to to VCs and to pension funds and the people. I think it’s I love Cohere. I love what we build and I also like talking. So I like talking about both
very similar.
Oh, between them? Yeah. Yeah, actually. That’s an interesting question. Yeah. I I do think look, the industry is a lot more matured. Two years ago, you know, when we were fundraising or three years ago, a lot of the questions were like, what is this? Like, what? What? How are you gonna what? How does that work? What is this? You know, and so we’d spend more time explaining stuff. People mostly know how it works and know what it does. And now we can say, here’s what we’re doing for our customers specifically.
You know, like here’s how RBC is using it. Here’s what we’re doing with Fujitsu. Here’s what LG is doing with it. You know, like we can talk about those things specifically. That’s more interesting. How much of the 600,000,000 will be spent on compute? Yeah. There’s like three components that go into making language models. Right? There’s talent. There’s like people. There’s some art engineers and researchers. There’s compute and there’s data. The importance of those has shifted and they yeah. The spend of those has shifted over time.
We train very efficiently. We train efficient models. So, like, our our model Command A and the Command A reasoning model, which we just released, those they’re all trained to fit on two GPUs. That’s, a really important part of our business strategy. It turns out if you go talk to a lot of companies who wanted to deploy models into production, they were bottlenecked on deploying because they don’t have enough GPUs. Two GPUs turns out to be like a sweet spot between performance and cost and, like, and actually how many GPUs they had access to.
So that means we train very efficiently as well. We have spent orders of magnitude less on creating foundational models than some of the other foundational model companies out there. Truly orders of magnitude less. And I’m very proud of the efficiency of the team and like what they’ve done with the, you know, the resources that they have. We think about efficiency a lot for ourselves and for our customers. And those two things are are related. But how much of our funding goes to compute? It shifts over the years, but a lot.
Compute is is a it’s How has it shifted over the years? I mean, when we first started Cohere, one one of the very first things we did, because we had no funding, was we we spent next to nothing on compute. And we showed that you could train a model by having like a bit of a GPU over here and a bit of GPU over here, bit of GPU over here, and you could link them together. And we published a few papers on that on training models with, like, the scraps of GPUs in data centers.
Right? That was what we started with. And we showed that you could do that. You can do that. It’s very slow and it’s much easier to just rent a big data center and train the model there.
The question that everyone asks is, how do you compete against competitors who have billions and billions of dollars? Do you hate that question? And how do you respond to that? No. I don’t hate that question.
I think yeah. I think that’s a fine question. We’ve we’ve announced funding rounds. They are smaller than some of the other funding rounds out We’re pretty singularly focused in a way that the other companies who build foundational models are not. Right? Like, we don’t have a consumer app. We’re not trying to get anybody to spend $200 a month on something for their personal lives. We’re singularly focused on working with enterprises and businesses, making sure that they get to production with AI. I have my whole I’m constantly telling people, like, not AGI, ROI.
ROI, not AGI. There’s a lot of work that still needs to get done there using something Do you
just think then that OpenAI and Anthropic will just seed enterprise?
I don’t know. I think like right now, you know, those both those companies have a pretty cool. They’ve both made good consumer products. Where this technology adds the most value is in work for, like, personal reasons. Like, that’s where I see this technology being the most useful. I don’t know if they’ll start working on that. I know that making models that work in that environment is pretty different than making a model that works environment. In a consumer environment, you can make the biggest model possible. You can have, like, complicated switches to tell you to go to this model or that model because you’re just posting it on a huge amount of GPUs.
You can be, like, losing a ton of money on every inference call, but, you know, you’re getting you’re getting users and something and so, like, that works. The types of models you have to build to succeed there are different. Another work you need to do on the interface, like, know, we’ve we’ve announced North, which is our agentic framework. It’s privately deployable, customizable for knowledge work workers within an enterprise. Is pretty diff it looks pretty different than some of the consumer applications. Right? Like a big one is like our models and generate images.
Nobody in the workforce is really wanting to generate images as part of their work. But as a consumer, it’s very fun. It’s very cool to be like, oh, here, give me a picture of this or something. So we, the types of models we train are different and the interfaces we make are different. I don’t know if that’s, if they’ll be interested in that at some point. I think, like, we stay focused on talking to customers and adding value. How do you price? Entirely dependent on what the customer wants to do with us.
So we do have some customers where we give, like make a custom model for them and give them that model. So do you have forward deployed engineers? We do, yeah. So they’re a crucial component of how we like go get a company up and running and into production with us.
Do you think everyone will have forward deployed engineers in a future AI world in a way that Palantir has glamorized? Yeah. I mean,
I think forward forward deployed engineers are a good idea. Right? Like, you’re selling technology to somebody. It makes sense to have some engineers who come and help them get it set up and work with them to like make sure it’s actually delivering value. You know, I think that’s a good idea. I don’t know if that’s true for every business. Does FTEs not just allow for poor technology? No. No. I think that there’s this there’s this idea sometimes like, yeah, you can just make the thing and for every business, it’ll work perfectly and require no engagement.
That’s the way some technology works. That’s the way a lot of consumer technology works. That’s not the way a lot of enterprise technology works. Right? Like there’s you’re selling things to people that have to be like matched to the way their business is set up. And so having engineers go along with it and say, cool, here, here’s here’s the model, here’s what we can do to make sure that that’s perfect for you in your specific use case is helpful.
Given that you sell to enterprises, I’m an enterprise investor and I love enterprise. Revenue quality is much higher, much stickier. Growth is slower because you’re working with large enterprises. Do you think you hadn’t and have an enterprise discount applied to valuation because of revenue growth being slower because of enterprise? It’s such a question. What was the price on the last round? It public, I think. It was, 6.7. Yeah. 6.8. 6.8. Yeah.
Yeah. It’s these are all staggering numbers. These are all numbers that are impossible for an individual to conceive of. I think we’re we’re we’re so far into that. You know, as a as a single individual, this is well beyond the realm of what you can reasonably engage with in your life. Is it? For a regular person who grew up working as a cook, yeah. Like, my first job was burgers. Right? Yeah. Yeah. That’s a great know, these are all crazy numbers. Do you care about money?
Yeah. I think yeah. Certainly. I think everybody cares about
money and everybody’s motivated by money. With being motivated by money, you’re an incredibly acquisitive asset, and it’s been a very strategically important thing for large players to do. Have you had M and A offers across the journey? Oh, yeah. We have at times. Yeah. How’s the decision making gone there? I always wanna be in the room. I always picture it kind of like thundery nights. And people coming together when it’s raining outside. No. No. No. No.
No. Yeah. We have. I mean, look, we’ve been a company for five years. Yeah. We have.
Not tempting.
We’re all being like the the co founders and now, you know, the people who work here, we’re all a lot of things you’ll hear Aidan say is like building a generational company. You know, we’re all really interested in building something that outlasts us and that goes goes beyond our involvement of it. That’s really exciting.
Why do you want that? I know it sounds strange. Why do want something that outlasts you?
Oh, yeah. That’s Remember earlier when I was like, sometimes I can ask philosophical questions and then we deviate too far off the thing and this is one of those questions.
No, it’s cool. I love this. Yeah. Is why the people say like, oh, AI could just replace you as an interviewer, Harry. It’s like, no, it couldn’t because it doesn’t have the ambiguity to go off on the like, but why does that actually matter?
Well, then if you think what you just said, why do you not think that that exists for all jobs?
Oh, because I think what I do is a very disciplined art honed over ten years compared to a social media manager who’s 24 coming out of university writing, we’re thrilled to publish our latest report.
That’s everybody thinks that what they do is is honed and trained.
And why is mining paid millions and theirs not? Because society places strategically more value on mine if we’re being a dick and blunt, and I’ll take this out, because very few people can do it.
There’s labor that’s easier to work, that you can learn faster, or, like, you know, you can get up to speed quicker. And there’s things that take a really long time to do. And, like, the only way you’re gonna be able to do that job is if you spend a really long time doing it. And there are some things that are harder and some things that are more agent like, that have more agency. And our economy is, you know, decent at figuring that out and compensating people based on on the on the investment that they had to make and the skills they had to have to get there.
But I don’t think it’s perfect. And I think everybody thinks and should think that the work that they do is a skill that they learned. And even work, like, some of the hardest jobs, some of the hardest days I ever had at work was, like, working at the grill, making breakfast and burgers for people when it was super chaotic and and stressful and there was no air conditioning because it had broken and, like, I had to run across the street to, you know, buy extra potatoes because we hadn’t prepared it.
Right? Like, that was challenging and rewarding work. And I don’t think just because I was getting paid minimum wage at the time means it it wasn’t valuable.
But it is it, like, it is definitively less valuable. It’s it’s definitively paid less. Adam Smith’s invisible hand would suggest that it is just definitively less valuable that you went and got, like, more potatoes, which meant more chips for someone that didn’t probably need more chips in a grill that was a single consumption model.
Yeah. I mean I’m sorry. No. No. No. Like, I I understand you’re work
that your team do, which which has impact on thousands of employees in some of the biggest companies in the world. Look, I’m glad I do the work I do now. I’m doing impact work on millions of Fujitsu users or Yeah. Yeah. It is legitimately, definitively less impactful and less valuable work.
Yeah. Again, like, as with a few times over this, like, there’s some extreme that says, hey, the invisible hand is absolutely accurate and, like, whatever you’re paid is exactly as much value you’re creating and exactly as much value as you’re worth. And then there’s another side of it that’s like, whatever, everything is the same. Who knows? Who has any idea? It’s all the same. Right? There is some middle ground. But
So going back, why does it matter that you have a generational company? After we detour it Yeah.
Seriously. Out of the spiffy of
Like, one thing that comes to mind is like, you know, look upon my works in despair, like, you know, Ozymandias and the idea of, like, people obsessing over their legacy and building, you know, you know, some statue to their grandeur. And, like, one day, it it will also fall. Right? Like, one day, all that will be left are two legs in the desert. That’s true. And that’s true regardless of what you build. But when I think about building it, like, excites me about building Cohere, and when I say, like, a generational company, I mean timescale generations.
I don’t mean, like, my generations. I mean, the idea of building something that is there for a long time. It’s rewarding. And I think it’s inherently human. You know, like, I think we all like to think about what are we building? How long is it gonna be there? And whether that’s like a work of art or an actual building or a company or a philosophy or an idea, the idea of building something or participating in the construction of something that is bigger than you is rewarding.
Totally. And it’s, like, fundamentally human. Even though at some point, yes, it will be two feet in the desert. Both those things are true. You know? It’s it’s rewarding and exciting and ultimately, you know, as with all things.
What was the biggest disagreement that you and Aidan have had? That’s also
a good question. The biggest disagreement. We had some different disagreements about, like, API design. Was a brief moment, like, before RLHF where we were talking about, like, oh, we should make an endpoint for summarization, an endpoint for entity extraction or something. And I think we disagreed about that. So we disagreed on, like, some low level stuff. But beyond that, you know, like, I’ve had the privilege of getting to work with both Aidan and Ivan, the other co founders, like, you know, I have a huge amount of respect for and we definitely disagree and argue about like the the little things about how to run the business.
Should we do this or should we make that policy? But there hasn’t been any like
I was chatting to Aravind from Perplexity a month or so ago. Yeah. I I actually chatting to him, it sounds so like behind the scenes. Yeah. Remember on stage in front of 4,000 people or whatever it was. It wasn’t a private conversation. And I said to him, it feels to me like you, Sam, Dario, all leaders of foundational model companies, you’re basically just kinda like presidents who sit on top of the machine kind of shouting views because we’re in a shouting views world and then everyone in the machine does the work.
And he was like, that’s exactly what we’re all doing. Yeah. Wow. Me, Sam, Dario, we just have to be, like, in front of every camera, doing every interview, basically espousing the views of the organization constantly because it is so important to be front and center and relevant today. Do you feel that Cohere is telling your story enough publicly? But they’re all consumer companies. They all fundamentally
make money via subscriptions from consumers. And the Anthropic quasi,
I mean, you’d say the majority is enterprise.
I would say a lot of it’s most, what is it like API calls from coders, right? Like, that’s a huge amount of their So, like, you know, maybe it’s like, I think, whatever, 50% Cursor or something. And now, yeah, competing with Cursor. But a lot of it still, like, comes down to an individual decision of a of a Yeah. So there, I think I understand their motivation. Like, yeah, if you’re if you’re selling to consumers, you wanna be telling a story, consumers are really interested in that story right now.
Yeah. That kinda makes sense for them. That’s not what we’re doing. You can’t you can’t spend $200 a month on Cohere as a person. Like, we we don’t have that offering. So it’s not as important for us. Like, could we be doing a better job, you know, telling what we’re doing and when I’m asked to come on here, yeah, I’m excited to come. I’m excited to talk to you. I like talking about Cohere. I think it’s important. Do I think it’s the most important thing? Like, no.
Building a product is the most important thing. I think solving problems for our customers is the most important thing. I think making a better model for them is the most important thing. Telling our story Maybe
that’s more important than discussing Adam Smith’s Invisible Hand with me. I do.
Thought I thought I did enjoy it. Yeah.
But, yeah, I do. That’s hilarious. One area that we haven’t covered, which is interesting, is the area of sovereignty. We’re sitting in London now. And Mistral’s in Paris, and we all say that for Mistral, it’s like the Europe play, and that’s why it’s funded and it’s continued to be funded. Do you think that we will see sovereign models and usage because of geography?
I think this technology is is a lot like infrastructure. Right? Like, I think building a like, having a language model that speaks the the language of your country is like building infrastructure for the people of your country. So I think that’s broadly a good idea. In the past, like, twenty years longer of technological history has been very defined by Silicon Valley. And I think a lot of people are not very happy about that. A lot of people are rightfully upset with some of the developments. Right?
Like, I used to be a real technological optimist. Like, I used to love the way technology was built and be like, oh, it’s so exciting or something. I I I wouldn’t describe myself as a technological optimist over the past ten years.
Wow. Why? What happened to change that?
Oh. Well, Sorry. Wait. Let me let me answer that. First, let me get back to the the sovereignty thing before I go off on this tangent. So, yeah, I think there’s a lot of people who are interested in building that infrastructure within their country and having the technology for their economies. Just using model that is built by China or built within America might not set your com your country and your economy up as well as having a model that it understands the context built in that language, in that dialect, in, like, know, has the has the cultural fluency needed to empower the people of the country.
So I think that’s like a good idea. What that ends up looking like, like, I don’t I’m not exactly sure.
Geopolitics obviously influences a lot. Do you think geopolitics has influenced customer decisions around sovereignty of models in the discussions that you see? I think us being Canadian is is an asset. That’s helpful for people. Do Canadian companies
wanna buy you more? Companies around the world are interested in talking to us. And in part, that’s because we’re Canadian. You know, over the past few years, you know, America has shown that they’re willing to, like, turn off access to tech based on political reasons. You know, we’ve seen connections between American tech and the American government is, like, less clear as time goes on.
What does that mean? It means, like, Trump influences US tech companies? Seems to be. Yeah.
Yeah. Right. I mean, even it was like last week or something, they announced they’re taking a 10% stake in Intel. Right? Like, that’s an interesting development. I I’m not an I’m not an economist. And I don’t know if that’s good for the country or not, but it is an interesting development. So I think there’s a lot of companies in Canada and around the world that are interested in working with non American tech companies. And I would say that’s been an asset.
Do you think governments should fund sovereign models? Is it a European imperative for us to have Mr. Al as an asset for Europe?
I think it’s a good idea for countries to have infrastructure within their countries. Like, I think it’s a good idea for people to have power plants in the country. You know, I like that Canada has several nuclear power plants and has several water power plants, like, you know, that that’s great. Language models are not that dissimilar from infrastructure.
Do you think our primary input device will still be a phone in five years time?
I do think language like, know language is gonna be a more important part of it. Like, I think fundamentally the way we should be interacting with computers is using language for the majority of it. Not all of it. There are times when, like, language is actually not the best way of interacting with a computer. It’s much better to have a graphic user interface where you’re, like, doing something. I know, like, last year there was a few there was, like, the RabbitR one. There was those, like, the humane pin, and, like, none of they didn’t really get it right.
But I think there was something cool there about, like, hey, how do we use a language model to work with a computer better? I haven’t seen it done right yet and I don’t know if it will. And I don’t know if that’s because going back to the technological optimism thing, like, you know, I was really excited when Google Glass came out. I thought that was really cool. Yeah. So was I. Yeah.
I I had it as a profile picture.
Yeah. Yeah. And then I got on a bus one time and somebody was wearing a Google Glass and they were delicate. They were like and suddenly everybody saw it immediately. You know, I was really excited about VR for a while. And then I realized I actually don’t wanna strap a computer to my face. I’m not interested in being disengaged from the world more. I wanna be engaged in the world more than I am. I don’t I don’t want more things removing me.
Do you just worry about this is so massive, like, this is the state of the world in terms of depression, in terms of loneliness? You know, the biggest pandemic, epidemic, whatever we wanna never know the difference between pandemic This and
is such a podcast. I haven’t done many podcasts. This is the most podcast podcast I’ve ever heard.
No. I just I get it. I can just do it. It’s too interesting where I’m like, fuck it. No, no, these are interesting questions I’m having. I’m just so nervous about the state of loneliness, insecurity, eating disorders, focus on materiality for young people. The number one job that any young person wants to be is an influencer.
Yeah. There are there are things that you’re talking about that I do worry about. I do worry about, yeah, the dissolution of community. And I think, like, what I’m talking about earlier was saying I wanna be engaged more in the world. I want the technology that I use to connect me to the world better. I don’t want it to disconnect me from the world. I think a lot of people feel that. I think a lot of people are looking for ways of connecting to technology more.
Right? Like, I I play music. A lot of the reason I play music is because it’s immediate. It connects you to the people who you’re playing with and the people who are listening. And a lot of the people who come listen to the music are there to connect in the moment. So I think a lot of people are feeling that. And I think a lot of people are feeling that because of the because they’re experiencing what you’re describing statistically in their personal lives. I also know that that worry of like, oh, no.
The world these days is, you know, so bad and things are going in the wrong direction and the kids these days are so weird and like, oh, no. Back you know, things used to be better back then, is historically ubiquitous. And everybody has always thought that going back to, like, Greek philosophers bemoaning the prevalence of writing because it was gonna, you know, make people not use their memories anymore. And saying, oh, no, the kids these days don’t understand honor and, like, going back to people belonging spread of newspapers because they were all sitting on the bus reading newspapers as opposed to sitting on the bus, like, talking to each other.
Like, I think two mutually exclusive things. One, yes, I’m worried about all the stuff you’re talking about, and I think technology and people who make technology need to think very hard about if their technology is helping with that or hurting that. Two, this everybody’s always thought that that that the time that they’re alive is the time when it’s the worst. And those you have to hold both of those two
conflicting views in your in your mind. We’ll do a quick fire, but you I’ll send you off this brilliant song, and it’s not really a song, it’s like a commencement speech by Baz Lerman and it’s called Wear Sunscreen. And it basically says exactly this, which is every generation always looks back and goes, oh, prices are so high today and kids are so rude today and it was better in my time. It’s this kind of continuous pattern of life in terms of looking back and thinking it was better than the one we have today.
Yeah. That that I mean, like talking about intrinsically human things, like that also seems to be intrinsically human, you know, wanting to be part of something bigger, wanting to build something that outlast you human,
thinking things were better when you were young. We’re gonna do a quick fire. If you were Sam Altman today, what would you be doing that he’s not doing?
I don’t think Sam Altman has done a service to the world by talking about how close AGI is. I think he has made several predictions now that are wrong and that were obviously wrong at the time he made them. Which one is most prescient? Oh, that like AI is gonna kill the whole world in two years. Like, he’s, you know, he’s made allusions to things like he did a world tour where he spoke to every major leader the world over to tell them, hey, this technology is gonna you know, poses an existential threat.
And I think that was academically disingenuous and I think did a disservice to the technology he loves.
Do you not see a correlation between the words that one has with regards to the future of AGI and AI and their requirements for funding?
I don’t I don’t know what that Do see what I mean by which is that statements and stuff
Yeah. Yeah. For a long time do not need funding. Yeah. And they are much more balanced, neutral. Yeah. And then other people who do need funding have to be much more provocative and out there because they need your fucking dollars.
Yeah. Yeah. I don’t know if that was the strategy. The correlation you’re pointing at exists. I would say that, you know, we’re a we’re a venture capital funded company and we need funding and we don’t say that.
Right? What worries me though actually is even the rhetoric from your and your has changed. Even their aggression towards the changes that are coming has has flipped, which does make me worry. For what reason? For the reason that actually we are far closer than we think to very material shifts in labor patterns, workforce behaviors. When even Zack, who does not need the money from anyone or Dennis, who doesn’t need it from anyone is going, shit. The changes are real.
Yeah. Well, there are some real changes. Right? Like, not don’t wanna you know, this technology, fundamentally transformative. The same way the personal computer fundamentally transformative, the industrial revolution, steam engines, the printing press. Those are all big technologies. There’s tons of legitimate things to talk about, and I’m glad people are talking about them. There’s also a whole lot of not legitimate things to talk about as people are spending their time on.
What is your founder ritual after closing each round? Who told you about that? Jordan. Okay.
That’s funny. We
go to
McDonald’s. You go to McDonald’s? We go to McDonald’s. Yeah.
And have the same meal? Yeah. I normally get two junior chickens. I don’t remember why that started. Wow. What’s the worst thing that can happen with regulation towards AI? I think the worst thing that
could happen is that out of an erroneous understanding of the technology and, you know, thinking that what we’re building is is digital gods, which is like a lot of large languages, models are not. But if you think that that’s what they’re building, then you could think, okay, cool. We need to, you know, come up with benchmarks around existential threats. There are times when looking at those at benchmarks, like, you know, fixation on particular benchmarks, which can be gamed and can be trained either to do way better on or way worse on, are not helpful for establishing how the technology can be used and misused.
So I think if you were, like, worst thing a regulation could do would say, hey, we’re gonna pick this random benchmark, we think that represents AGI and we’re gonna shut down any development on it. I think that would be that would be a misplay.
Do you think China will produce leading models in the next two years that continue to beat US models?
No. I’m not sure. They haven’t yet. They’ve made good models. Definitely good models. But I don’t think they’ve made models that are like beating, you know, the other models out there.
You’re not worried by China in their model capabilities. When I look at cadence of China, in like a week, about a month ago, there were like seven new model providers, really seven models. They were pretty good.
Yeah. I don’t think I’m like worried about it. I think they’re gonna keep building models and those models will be useful and they’ll be particularly good and the things that they’ve trained them on. But that doesn’t cause me like a bunch of What’s your boldest prediction for LLMs in 2026? In 2026, you’ll be able to open up a computer, log into North or whatever application you’re using and say, file my expenses. And then the model will, yeah, figure out what expense policy it is and what the where the photos are, like, do all that.
Like, that’s do all of that for you. That’s my boldest prediction. And I know that’s not very bold. I know in some ways that’s like, oh, that sounds like not too far off. And you’re getting it to actually work and be a thing you can rely on is is not in every company. Most people don’t have the experience I just described. And I think that becoming a ubiquitous way of using a computer is crazy.
If you weren’t at Cohere, which AI company would you bet your career on? Google’s great. Google DeepMind is building cool stuff, you know. That’s exciting. Have there been any tools that you’ve added to your workflow that meaningfully increase your productivity? So like for me, like Whisper Flow.
Oh, Whisper. Yeah. Yeah. Cursor. Cursor’s a great coding application. Do you
mandate Cursor across the whole engine team? No. Not at all. Does anyone use Windsurf or Devon? I think so. I’m not sure. We don’t mandate one. I know lots of people rely on models. Are you price sensitive towards Cursor costs increasing?
No,
but
I live a life of privilege.
So like, no, I’m not as a person. But when cost goes up 10x and you have a 100 engineers? Oh, as a business. Yes. As a business, obviously. How does that shake out? Do
we see this kind of like For us, yeah. I mean, we make our own models. So if we wanted to use our own model and use, plug that into an extension of Versus Code, would be something we could do if that was Would the quality of output be similar? Right now, no, no, no, no. Cursor has built a good product, like they built, know, they’ve done really good UX stuff and it’s cool. But if it was to go up 1,000 fold, like, yeah, of course, then we’d start thinking that.
Will we see a trillion dollar AI company outside of The US in the next decade?
Yeah.
Maybe north of the border. Okay. There’s one other one that I love, which is like, what trait do you love and has contributed to a lot of your success, but you’re also quite wary of? Of me?
Oh. Oh, that’s much easier to answer. I’m quite curious and contrarian, and that is both an asset and a hindrance. There are times when it’s very helpful to be like, oh, yeah, I’m like super interested in something and I learned about it and I’m like, everybody thinks this and they’re totally wrong. That’s super helpful. That’s why when Aidan was like, hey, do you wanna found a company on language models back in 2019? I was like, yeah, absolutely. You know, that was not a view that was widespread.
So I’m, yeah, curious and contrarian. But there’s other times when like the whole world has been definitely right and I’ve been wrong And I’ve been like, oh, yeah, was super excited about this. When were
you most wrong? I’m most wrong.
When I was very young, I was very much a, like, a technological optimist, as mentioned. But I felt like, cool. All all metrics of human improvement are gonna are gonna continue to go up. You know, like, life expectancy will go up. Income inequality will go down. Happiness will go up. Like, as as yeah. We’re just the the path towards humans endeavors and success is monotonic. And that’s not true. You know, there are times like that. That was that was something I was really wrong about. Yeah.
Another thing I was wrong about was the efficiency, the data efficiency of reinforcement learning from human feedback. That was just to give a real technical answer. This is back in, like, 2020. And I remember being like, oh, no. You can’t you can’t make a small dataset of feedback from people and make a model better. That was a that was a technological misstep.
Nick, this has been a a a interview of many twists and turns. It’s the joys of doing what I do that actually, which is like, the natural conversation that’s inspired. So thank you so much for agreeing to partake in such a wide ranging conversation. Thanks for
thanks for having me.
I wanna say a huge thank you to Nick for joining me in the studio. And if you wanna watch the episode in full, you can find it in video on YouTube by searching for 20 VC. That’s two zero VC on YouTube. But before we leave you today,
· Sponsor read0 min · 627 words
I love seeing the team come together to make this show happen. What I don’t love is trying to keep track of all the information, the data, and the projects that we’re working on across dozens of platforms, products, and tools. That’s why we use Coda, the all in one collaborative workspace that’s helped 50,000 teams all over the world get on the same page. Offering flexibility of docs with the structure of spreadsheets, Coda facilitates deeper teamwork and quicker creativity. And their turnkey AI solution, the intelligence of Coda Brain, is a game changer.
Powered by Grammarly, Coda is entering a new phase of innovation and expansion aiming to redefine productivity for the AI era. Whether you’re a startup looking to organize the chaos while staying nimble or an enterprise organization looking for better alignment, Coda matches your working style. Its seamless workspace connects to hundreds of your favorite tools, including Salesforce, Jira, Asana, and Figma, helping your teams transform their rituals and do more faster. Head over to coda.io/20vc right now and get six months off the team plan for startups for free.
That’s coda, coda,.io/20vc, and get six months off the team plan for free, coda.io/20vc. And while coda keeps our team aligned, let’s talk about Brex, the ultimate financial stack for startups. So when Brex was founded, it wasn’t just about creating another financial product. It was about solving the really gritty challenges that founders face daily. Let’s be honest. Building something from the ground up is hard enough without dealing with clunky outdated banks that pile on fees and leave your cash idle. Brex is different. It’s the financial stack that scales with you no matter where you are in your journey.
From corporate cards to maximizing your runway to earning yield on your cash. Brex was designed with founders in mind to make every dollar go further so you can focus on building. And here’s what really stands out to me. Brex combines the best of checking, treasury, and FDIC insurance in one powerhouse account. You can send and receive money globally at lightning speed, earn yield from day one, and still access your funds whenever you need. Plus, with 20 x the standard protection through program banks, your cash is not just working harder, it’s working safer too.
It’s no surprise that one in three venture backed startups in The US with companies like Anthropic, Coinbase, Robinhood. I mean, my god, these companies are incredible. Trust Brex to help them grow. If you wanna join the smartest startups on the planet, head over to brex.com/startups and see what they can do for you. And speaking of incredible companies, don’t forget what really keeps those customers coming back. Trust is the ultimate currency in business, and today customers expect it faster than ever. And that’s why over 10,000 global companies trust Vanta.
Vanta automates up to 90% of the work for in demand compliance standards like SOC two, ISO 27,001, and more using smart AI to centralize workflows, manage risk, and get you audit ready in weeks, not months. So you can stop chasing paperwork and start closing deals. And a new IDC report found that Vanta customers achieve $530,000 per year in benefits. That’s insane. And the platform pays for itself in three months. I had no idea about these. Whether you’re growing fast or just getting started, Vanta connects you with trusted auditors and experts, support to help you build trust with customers.
Get a thousand dollars off your first year at vanta.com/20vc. That’s vanta.com/20vc. As always, I so appreciate all your support, and stay tuned for an incredible episode on Thursday with Jason Lemkin, Rory O’Driscoll, and a special guest in the form of Canva cofounder, Cliff Obrecht.