Go back

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

47m 22s

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Chai Discovery, co-founded by Josh and Matt, is a foundation model lab for biology, aiming to transform drug discovery into a more engineering-like process. The core idea is to replace serendipitous trial-and-error with AI-driven design, where models can generate molecules with desired properties, similar to how code generation works. Key milestones include the advent of deep learning for protein folding (2018-2020), which evolved into generative approaches like diffusion models, allowing simultaneous generation of protein structures and sequences. This shift enabled tackling harder problems like antibody design, previously deemed infeasible due to data scarcity. The team, initially composed of AI researchers, has expanded to include engineers and antibody specialists as models improved. They emphasize simplicity in model architecture and high-quality software, with a small, over-capacity team to prioritize impactful work. Recent models achieve higher binding rates (improving from ~0.1% to more), moving toward zero-shot molecule design. This could increase lab testing ROI, paradoxically boosting demand for experimentation. Ultimately, Chai aims to industrialize drug design, making it more iterative and design-oriented, with the lab serving as verification rather than discovery.

Transcription

10704 Words, 57602 Characters

English
One of our big guiding principles is just simplicity. So when you look at a model, let's say a Chai-1, I think there are 23 distinct submodules in Chai-1. And when you're trying to iterate on something like that, it gets really hard. Because you're like, I need to understand each of these submodules independently. I need to understand all of their behaviors, their dynamics. And that doesn't actually scale that well. So you can start to think, OK, how do I simplify this? How do I identify what's really important? And once you have that, the whole research process and like, identifying these types of scaling directions becomes a lot simpler. [MUSIC PLAYING] We're here with Josh and Matt, two of the co-founders of Chai Discovery. Chai is engineering molecules with AI. So it's something of a foundation model lab for biology. Sounds like a big idea. But let me just start there. What is the big idea? Thanks for having us on the show. One of the exciting things we're trying to do at Chai is to make the drug discovery process look a little bit more like engineering. And we've seen all the work happening with L-Lums for code generation, for instance. And that works really well because code's a very simple abstraction to kind of get across what you want to do. Biology doesn't look like that today. It's a lot of trial and error. And I think if we go back to the early days of just like modern biotech, right? A lot of the medicines that we have today were kind of discovered quite randomly. It was a bit serendipitous. And what we're trying to do is allow us to industrialize that process a little bit more and try to come up with the tools that we'd need in order to have abstraction layers. The same way we have that in code and in modern engineering laws iterate really quickly and try to bring that into biology and into drug discovery. So let me ask you a question on that then. So tell me if this is a reasonable way to frame it. There's almost this boundary between that which can be engineered and that which needs to be tested in the real world. And it feels like that boundary has been moving over time, like the proportion of the drug development process that can be engineered as opposed to trial and error seems to be increasing. Is that a reasonable way to think about it? If so, can you talk about what specific developments have pushed that boundary over time? I think that's one way to think about it. The way that we often look at this is like the lab is an important part to verify that what you're doing is correct. And actually, verification is a very big theme in AI as well. If you can evaluate that your model works, you can verify it, then you can start to hill climb that and you can make progress on it. So I think the lab is a very important part of this. And the question is how do we take drug discovery and make it look a lot more like drug design? So one of the reasons why we call this field drug discovery is we're often looking for needle and a haystack. We'll screen millions, billions of molecules try to find one that works. And if we can instead put in what is the dream state of the molecule you want? And then the model can materialize that. That's going to be really powerful. So it's not even a question of reducing lab testing. I mean, maybe that happens as a result. I actually might even take the opposite side of the coin and we can talk about that where maybe we'd actually do even more lab testing because the ROI will increase. The same way there's more software engineers or there's more demand for software engineers. Now that they become more productive, there might be more demand for the lab. But I think the key part is how do we just change the paradigm here and how do we make it more design oriented? Can we talk about what is state of the art today? And then maybe let's take a little trip down memory lane. Five years ago, three years ago, one year ago, what have been some of the major breakthroughs and how has stated the art changed over the recent years? Yeah, I think the field has-- it's really evolved, especially over the last decade. So it really wasn't until you can actually look back. There's this biannual protein folding competition. So every two years, usually a bunch of academic groups who compete in this protein folding competition. What would happen is you hold out these protein targets that no one's ever seen before, they never get deposited publicly, and then all these groups compete so you can predict the proteins the best. And really, it wasn't until 2018 where you started to see this big step change in performance. Then finally, again, in 2020 with Out Fold 2. But it was kind of like the advent of deep learning. That really kicked off this whole field. So at first, what you do, these are kind of like-- bring me back to my original days and my advisor's lab. So we worked on one of the first systems to do this with deep learning. What we would try to do is predict the distances between amino acids and a protein. And then we'd have some render, which then took in these kind of noisy and complete looking distances, and then emitted some protein from that. And what's pretty remarkable is that with relatively minimal information, you could build a machine learning system that could actually produce something that really looks like a protein. And for the most part was correct. There were still pretty big gaps. It didn't quite have the resolution of what we have today, almost like with the early image generation models, where it was a little bit grainy, a little bit blurry. And now it's like, oh my gosh, this is crazy. High definition images in seconds. So we kind of saw that same evolution happen over time. So it started with protein folding. And then I guess once that became more realizable, there were kind of these other sub-problems that people wanted to solve. And now given a protein structure, can I design say a sequence which might fold to that? And that's getting closer and closer to the drug design problem. So you're now thinking, OK, how do I start engineering proteins that match a certain shape or might perform a specific function? And so people kind of studied that problem independently, and that was kind of like 2021, 2022 time. And then these ideas began to merge together. It was really like the advent of diffusion models, where we started to be able to-- OK, I now generate a protein structure in a sequence kind of simultaneously. And I can start making these like what we'd call prompts more and more realistic. So I can now prompt these models on a target that might have a certain shape. And I can say, hey, I want to bind over here, kind of like add more real world constraints on the problem. What did you guys see in 2024 that made you think that was the right time to start the company? Yeah, there were a lot of discussions that went into it. I remember one of these early discussions actually back in Matt's house when we were-- Matt was showing me some of these results on like antibody antigen, like structure prediction. And Matt was just talking about protein folding. But for a long time, people thought that protein folding for antibodies was just like too hard of a problem. People were like, there was not enough data for antibodies like in the protein data bank, for instance, to solve this problem. And people as a result thought that like antibody design was going to be out of reach. A lot of the work that people even did on protein design in the early days with deep learning. It wasn't antibodies. It was these different class of proteins called mini proteins, which are actually really interesting in their own right. But they're not what most of the drug industry is looking at. So the Holy Grail was whether we could design antibody proteins on the computer, especially ones that had all that therapeutic function. And our thinking was that if you couldn't predict what an antibody looks like, how are you ever going to design one? Traditionally, people have thought like biology is scary. Like there's so much to know-- I'm terrified of biology. I am as well. Honestly, my background was never biology. I studied like pure math and started my PhD in theoretical computer science. And it was only like after my third year that I ended up switching into like deep learning, protein, structure prediction, all of this stuff. So it was like totally new field to me, seemed insane. But like at the end of the day, it's much simpler. And like the problems are much more interconnected than one might think. So like I think people like when they start going into field, they're like, oh man, what's an antibody? What's a mini protein? These are all just like sequences of amino acids. At the end of the day, like these are just like different types of prompts for the model. But like in the same way, we might like have a math problem that goes into chat GBT. Chat GBT can both answer your math problem and like help you with your English homework. So like really, we have the same type of thing going on with our models. Like we just have some way of representing these sequences of amino acids. Then we have a way of like designing, predicting those as well. And I think like in that lens, like things become a lot more clear. And Matt, you mentioned your background a little bit. Josh, talk a bit about your background and then more broadly. In order to pull this off, you have a bunch of different disciplines that kind of come together. So can you just talk a bit about like your background and that of some of the other core members of the team and how these things all fit together? Yeah, I've been excited about AI and biology since I was a kid. So I guess I'm like Matt, I didn't start with like theoretical CS and get into that way. But I went to a high school with a STEM cell lab. So I was just always excited about biology as a kid and a group as a programmer. I really started my career at OpenAI, so it was on the early team there. It was a nonprofit back then. So it was a pretty good time to be there. We did GPT1, GPT2, scaling laws. And the question was like, if the models can learn to speak English, German, French, why can't they learn to speak DNA and protein? And that was kind of my research agenda since then. And I think that sort of intersects with around the time like Matt got into the field as well. And it was a, I think it was a pretty important time as well, right? If you look at the kind of methods that we were bringing in, there's been a lot of these changes on the edges, if you will, right? And Matt talked a bit about the history of what's happened in the fields here. But it's all about how do we find the right deep learning architectures with the right compute configuration and the right model architectures to make this happen? What are the right tasks to apply it to? We're talking about how we even knew that like, 2024 is the right time to start the company. As Matt was saying, these are all different kinds of amino acid sequences. And people thought that, you know, anybody class of problem was going to be too hard. And we started to see the first signs of life that actually this was starting to work. I think a lot of it fueled by some of the new architectures we're bringing to the problems, right? I think it was diffusion models back then. The first time like anyone was able to generate reasonable looking proteins was the advent of diffusion models. So it was pretty crazy. like they were a bunch of generative modeling approaches that like would kind of work if you if you had a bunch of data So like people got these working for images There are like a bunch of tricks to make this better and better along the way Like we've definitely borrowed a lot of those ideas in our domain as well But it was really like once diffusion models came around. Is there an intuition for why diffusion models work? Yeah, so diffusion is is not a once that process so like I think kind of up until this point like the the main Generative design paradigm was like called variational auto encoders And in that case you're like saying I want to just like compress my input distribution So like you you have some like proteins you want to make these look like kind of like fuzzy Gaussian vectors That's has to just like really hard and maybe today if we tried like super hard I think we crack it But at the time like I think like it wasn't really as developed enough to work for our problems What turned out working really well It's just kind of giving the model more time to think and showing it like more examples like here's like a slightly broken Looking protein. How do you make it better? You can kind of break that protein more and more and more and you can make it look more and more noisy more and more broken And teach the model just quick little shortcuts. All right. Here's how I make it slightly better And you can just keep asking over and over again make it slightly better make it slightly better and kind of like Breaking the problem down to that scale worked really really well for for biology Yeah, the make it slightly better reminds me of a game that I'd like to play with my daughters Where we have chat GPT give us a unicorn and then we make it stronger and we just keep telling it to make it stronger And by the time we're done we have the strongest unicorn in the world. So about the same, right? Yeah, that's It's that easy. Okay, maybe not the same. Yeah, so I feel like in this domain you need to assemble a kind of like a quadrary Lingle group of people like an Avengers Squad of chemistry people biologists AI folks and so that's a challenge How have you guys gone about finding people convincing people to join the team and who are your superstars? So we've been really pragmatic about this such a if you look at the founding team It was mostly AI researchers So people who had Worked on either scaling models or getting them to work in this domain But really with each generation of model the kind of people we've needed for the next milestone has has changed or I'd say probably as expanded right So if you look at try to right that's the point when we started to bring in Some of the most incredible like you know anybody engineers and scientists in the world One of the son of stream Andy and a young actually when we hired him People asked us if we had pivoted into building a full-stack drug pipeline because they're like you'd be crazy not to do that if like Andy joined your team But I think Andy has has enjoyed running more antibody campaigns in the past couple of months Then he's probably running his whole career, which is very cool to see you have folks like Nathan Rowlands on the team like Nathan was actually homeschooled and then started college very early on so you're like David Daikers lab who won the Nobel Prize for protein design like when he was 14 Stars PhD when he was 18 and has so many creative ideas on our as the model started to get better We needed to build up a product team because while the researchers might get the models to point that they're They're very powerful. You need to build the right products interfaces that the models are actually useful and that's where we started to bring on People who have you know built some of the most exciting products we know about today like our co-founder Jack worked at Stripe Monaz who is one of the top 10 code contributors at Stripe Neil who ran his own cybersecurity company before security started to become very important as we deployed this to our big partners as We started to scale up we brought in people who've Really done a lot of like the GPU hacking if you will in order to like scale up our systems that they they don't break when we're running them at scale We had an email from or a slack message from one of our hyper scalers the other day where we had a cluster I think that's a mischew to it and like the GPU's got too hot I think you guys are running too many and we're like isn't that the point right? That's probably I was like good. We're doing our job at least Yeah, it's like how do you convince those people to join? I think again a lot of it comes back to The results and like a clear need we didn't hire antibody engineers before and in antibody design model like what are those folks gonna do We were even I think worried when we started that trend because for some of the Next generation formats the complicated antibodies like multi specifics like they didn't even work with shy to So it actually took a couple of of weeks when some of those people showed up before the models could work at a point that they could work on some of these interesting case studies But fortunately progress is fast enough to kind of bring that online So I think we're always like evolving that team and going for you know that next milestone We've got the team very small as a result too So this way, you know everyone is a little bit like slightly over capacity I think which means we have to prioritize it forces us to work on the things that really matter most I think on the research side as well One of the founding engineers Kevin Wu he had the first I think was the first Protein diffusion model like ever and that speaks the Kevin's speed of execution like he is a heck of an engineer and like I think engineering Has just always been like important since day one So like really even our researchers are all excellent engineers and like we really care about building a high quality Co-base like at the end of the day we are technically like a software company We're AI researchers where pro-we're protein designers where we're a lot of things But our deliverable is some piece of software so we've kept that like kept that bar really high Well also trying to like, you know level that with great research talent and people that can actually like push the frontiers What's possible and you've had a number of it seems like aha moments in the field like alpha fold was an aha We can figure out how protein folds and then the diffusion models aha Like we can generate proteins And it seems like your latest models are a kind of another aha moment We're not only generating molecules that look like proteins, but they also have therapeutic properties like they can find really tightly They've really high hit rates. Can you talk sort of about the quality of the molecules that your models are producing now and everything that sort of went in To those models to make them able to do that Yeah, if we look at the the quality of the molecule that's coming out if it goes back to you know One of the theses and we started the company that we really wanted to focus in on this like denovode generation of the molecules If you look at what a lot of the drug discovery AI work was at the time is about how do I take a molecule Just make it a little bit better right which we just talked about in a sense But it was doing it with a lab in the loop style right where I take some data trying to make it better that way and The the question We started with is well can we actually just like do all that on the computer right is there a way that the there Maybe there would be enough data or we could collect enough data So that we could just zero shot a molecule that has a lot of these properties So the first thing we needed to do to get there was to design molecules with really high success rates When we started the company the state of the art for anybody design was about like a 0.1% binding rate So one and a thousand of the molecules you design would actually bind in the lab So first of all, I mean to have to screen a lot of molecules to find some good ones It also means that like the gradient you get on your process is actually quite weak as well So for many targets you won't get any hits for the ones that you know You do get some hits you won't have enough to actually see whether you're having like the drug-like properties So we really focused in on how do we just make this process more accurate? We got to with our chide 2 model about like a 15% success rate So now if you screen a thousand molecules you're getting 150 back Now you can get start to get some like interesting statistics on the properties of the molecules Right and a lot of a lot is to like iterate on that build the right evaluations around that in the lab and really try to He'll climb that as well and we're getting to a point now where we can actually bake in a lot of these different properties From the start into this engine and then maybe if you think about you know How does this happen or like how are we approaching this as well? I'm like why we think it's gonna continue improving. Yeah, so like Honestly, I think Josh and I were both surprised at how quickly this worked Like when we were originally budging this we're like ah maybe like 20% hit rate like three or four years We were like really shooting for like a 1% or we thought 1 and 100 would be amazing We're like this is gonna be a groundbreaking thing. We had a philosophy on like an approach that we wanted to take and it just like Ended up working really well and that approach is like very similar to what's worked in the rest machine learning So people kind of treat biology as this like bespoke problem or like bespoke field But really it's like the same principles is like self-driving LMS Well, some principles but one of the things we've talked about before is how you manage to find scaling laws And it's one thing to tokenize a string of text. It's another thing to tokenize biology Can you say a couple words about Without giving away any of the magic, you know, can you just say a couple words about that challenge and how you guys solve that? Yeah, so one thing that I really like about chai is like we're we're a very bitter less impaled company So like we really believe in like scaling data a scaling model scaling compute in order to do that Obviously you need to identify scaling laws otherwise you're just kind of like wasting time and resources And I think like without giving away too much like one of our big guiding principles is just simplicity So like when you look at a model like let's say a chai one So like they're like I think they're 23 distinct sub modules in chai one And like when you're trying to iterate on something like that it gets really hard Because you're like I kind of need to understand each of these sub modules independently I need to understand all their behaviors their dynamics and like that doesn't actually scale that well So like you can start to think okay, how do I simplify this? How to identify like what's what's really important? And once you have that like kind of the whole research process and like identifying these types of scaling directions Becomes a lot simpler one of the other things too I think it's interesting about this is we take a lot of these lessons from what's worked in the rest of the deep learning space But as you're pointing out like the data itself is different Yeah, the models that chai are completely built from scratch We're not like fine tuning GLM or something like that on some of the team data We build everything from the ground up I think a lot of the company building process though is is taking a philosophy and actually sticking with it and iterating on that and just having some guiding principles When you're building like a blue sky research company fuel, you know chai is almost like one of these neo-labs right like we have this big AI Problem we're going after where as we make progress on it, you know that opens up opportunity for our customers But if you're gonna work on something so open-ended that way you need some principles to guide you and I think we've done a very good job on tracking those principles in the company, working on it. So a lot of the things that-- - They can share them. What are the getting principles? - So I think simplicity was one of the ones that Matt mentioned. It's this bitter lesson, pillness of scaling compute and data and models. It's being really rigorous. That's something that's so important in this space. You can fool yourself so easily in biology. The error bars in the wet lab are actually quite large as well. So it's actually a little bit different than if you look at cogeneration, for instance. If you look at sweet bench people were like, oh, maybe a year ago, people were like, oh, I got 16% than 17% than 18%. In biology, if you're plus or minus 5% in your lab, that might all be the same. So it actually just means the bar is really high in terms of the step changes that you want to see with the models. But you also need to be really honest with yourself about whether you're making progress or not. So you could come up with some fancy model that looks like it works well in one or two new tasks. But it's very important to show that that works more generally if you're actually trying to build a product that can bring the field forward. Yeah, and that's I think pretty interesting, because biology is one of those inherently not-so-verifiable domains. And you guys have been really good at showing your progress to customers and to people like us, you know, very little about biology. And so could we talk a little bit about the evals and the verifiable part of the model progress? Like how do you guys know that your models are getting better? Well, I would say first of all that I'm actually think this is one of the more verifiable domains. It's actually very objective. Readout, if you look at something even like code gen, right? Like maybe as a verifiable task like did my code compile? Did it solve these unit tests? But how do you think about the taste, right? Like did I write some really sloppy code that can't be maintained? Like what does that look like? When we think about designing a molecule in the lab, we can actually be quite specific about many of these properties, right? So maybe we get a molecule that binds the target, but can I manufacture it, right? That might be your version of like some tech debt, but you can measure that. And I think those evals again actually make this domain more verifiable. Maybe it takes a little bit longer to validate it, right? It's not like five seconds to like get a readout and run a unit test. You might have to spend a couple of days, a couple of weeks in the lab to get that readout. But at least you can be honest with yourself. Yeah. When we talk about progress and how good the models are, there's like a domain of targets in biology that you can, as you mentioned, sort of address with traditional screening methods. They take a long time. They're very slow and rudimentary. And then there are targets that just aren't addressable with existing methods. They're not drugable for whatever reason. And so where are we in terms of model progress, in terms of working on existing targets and generating molecules faster? That's one end of the spectrum. And the other end of the spectrum is unlocking novel biology and new targets and things that we couldn't drug before. This kind of goes back to our core modeling philosophy. So there have actually been plenty of times where like, man, if we had a 24th module, we can actually unlock that new target. And we're like, is that really something that we want to maintain long term? Is this incremental or is this actually a compounding improvement? Well, this actually helps us generalize to the broader class of this whole class of targets that we really can't hit. And so our philosophy has been, all right, let's just continue to focus, identify your scaling laws. There comes a point where if the model is able to push lost down even further, it has to understand something very intrinsic about the target that it's operating on. So one example we were talking about the other day, Paul and I, maybe the way that we're looking at certain glycosylation sites on proteins, we're like, oh, we might want to represent them differently or something like this. And we're like, well, even if we dinner represent them, like there are certain sequence mentees that will tell the model, there should be a glycosylation site here. And in order to drive like lost down further, the model should just have to learn that. So there are all these hidden features of targets where if you really believe in scaling laws, you believe the models will get there, these types of targets should just unlock with better models. And of course, you still have to take this very seriously. And you still need all the proper validation. You need to really challenge yourself and make sure that this is truly working. But I think our approach has always been with better models. We should be able to unlock a lot of these targets. One of the other interesting things is if we look at-- we talked about the interdisciplinary nature of this-- if we look at the different teams at CHI, what people will call hard target is actually different in literally every team. So on the science team, it might be like a undruggable GPCR target known as God and something that has modulated that in a functional way. On the ML research team, it'll be something like, oh, the model just can't fold this thing up. It doesn't know what it looks like. And then on the product team, it might look like, oh, I've got all these modifications, my glycosylations, and it's a membrane protein. How do I represent that to the user? And actually, the fact that it's different for each of these groups is a feature rather than a bug. And it means that if we want to make broad progress over here, everyone is pushing and parallel on these different ways. And that means that there is very smooth progress that we can make all the time. And there's usually not one bottleneck at CHI. It's not like, oh, if we only had that one extra module all things would work. Or if we only try to push this into the product in some way, we could unlock it. We're trying to build this unified solution. Because on the day, the goal of companies to build a computer-aided design suite for molecules, it's not to make one or two molecules. It's not to get a pipeline of five interesting therapies that we bring to market. It's the change of the way that medicines are discovered. And we're going to do that. We need to work on all the hard targets regardless of how you define hard. Let's see more about this idea of computer-aided design suite for molecules. What does that mean? So at its core, it goes back to this point about making biology look more like an engineering discipline. So we're not going in fishing something out of a large library or doing a ton of trial and error. You want to be able to specify upfront the principles that go into designing your molecule. And then I have an engine that can actually realize that into some molecule that we're going to go and create in the lab. And look, you still might do some iteration on the lab and on your model because maybe your hypothesis was wrong. But what we want to speed up is actually, again, have that computer-aided design suite so that you can go from idea to testable hypothesis very quickly. And if that loop right now takes something like nine months, I don't know, to go and discover your molecule versus if it takes nine weeks or takes nine days. Each order of magnitude just scales in a very big way the number of ideas you can really sort through. I think that's ultimately how the field is going to converge on better medicines. It really comes back to people sometimes talk about, do we care-- and you kind of note it on it's not like, do we care about speed or do we care about the difficulty of the targets? At some point they converge in this way as well. Because a hard target, if we can iterate through hypotheses a lot faster, then maybe it'll be easier to crack it. So if we can maybe detach ourselves from the present reality and go far enough into the future that we're not thinking present-forward, we're actually thinking future back, 2035, 2040, 2100, whatever you want. What does the industry look like? Let's imagine that computer-aided design suite from molecules has become a standard. Let's imagine a lot of innovation has flowed downstream into some of the wet lab parts of the process. Can you paint the picture of what the industry might look like 10 years from now? Yeah, I think it's going to be a really exciting time. And we can look at this on a couple of angles. So first of all, the quality of the medicines that we develop will hopefully go up. There's a lot of molecules we put into the clinic today that there is really hard to discover molecule. You find something that's like 80% of the way there. Maybe I'm just going to advance it anyways. So I can hit my timelines. It's probably going to benefit some patients. But then you get beat a year later by someone else. And it's really just not the most efficient spend of resources in the industry. You've got the kind of diseases that are just too hard to go after today. People have been trying to drug Alzheimer's forever. And unfortunately, having made as much progress as we'd like. You have things that just aren't economical to go after today. Think about personalized medicines, rear diseases, things where maybe the patient population is going to be smaller. But again, if we can iterate through those ideas faster, if we can launch something faster, if we can do it in an unless expensive way, then those probably come into reach as well. So I think there are just so many different axes that we're able to push on. And I think that means that the future is really bright. It used to be like, you either want to be first in class or best in class. Yeah. And now it's like, you want to be last in class. Because you actually just want to be the final answer. So I think a lot of what you'll see is just way more intentionality in the types of drugs that you're designing. This drug will be super specific to the disease of interest. It won't have certain interactions, certain negative interactions that a lot of drugs today do. A lot of this stuff is actually like, you're able to model most of this computationally, maybe not today. But there's definitely a path towards getting there. And I think that's one of the most exciting things for me. Very cool. Can we talk about a business model decision that you guys made? Because I think a lot of times folks think about Chi and isomorphic and the same neighborhood. Isomorphic is developing drugs. You guys are enabling the existing industry to develop drugs more efficiently, better faster, cheaper than they have before. Why did you decide to go down that path versus the isomorphic path? Yeah. First of all, I think both of these paths are great. And they can create tremendous value. We've always been really excited about building infrastructure for the industry. Our bet when we started to go back to those results in Matt's house years ago was that this is how most future drugs are going to be discovered. And if that's the case, somebody needs to go and build that infrastructure to make it happen. I think part of this too is a lot of our founding team, including Matt and I had worked in companies before where we had built these full stack drug pipelines, where we were building AI models. We were putting those drugs into the clinic. Again, we're pretty bitter less than piled as well. And our thinking was, as the models get better, we want to be spending more of the money making better models as opposed to diverting those resources into clinical trials and things like that. And one of the interesting things about our business models as the models get better actually wins us the right to continue investing more in them. And you have partnerships with these pharma companies that are paying off today and allow you to continue investing it. So it's a much more scalable business in that way. And I just think about the alternate. an impact that we can create for the world is a lot larger. We've always wanted to just partner broadly with the ecosystem. It goes back to this point about fooling yourself in biology. If you work on a small number of drug targets, you might come up with the most exquisite molecules creating a ton of value for the world by doing that. But you might miss the force for the trees because maybe your model doesn't generalize to the other 500 molecules that people are going to make that year. And if you go and partner with people, you just can't fool yourself. Like, if you look at our partners, Eli Lilly, Novartis, Argenic, Pfizer, like these are not companies that are, you know, they take this stuff for granted. Like, you have to really deliver on these partnerships for them to take you seriously. And that means it can't just work like some of the time. Like when we ship models at Chai, they really have to work. They have to deliver value to our partners. And it's not like, oh, we made some molecule. It doesn't fully work. We're going to have our chemists like clean it up a little bit. So it's almost a harder business to pull off. I think that's one of the reasons why you haven't seen it pursued many times. If you work on a drug pipeline, again, like there might be ways to fix things up later. There will be the proof of like what happens in the clinic, of course. But when you have this partnering based model, you have to be really rigorous about your models. Your models have to work really well because otherwise those partners are not going to come easily. So it's made our life harder, I think, in many ways. But I think it's also the more rewarding path if we can get it to work. >> One, what have you learned? I mean, you know, you're out of the lab, so to speak. And in the real world, you know, delivering real value for actual pharma companies, what have you learned as you start to work with these partners in terms of any surprises in terms of how their needs might have differed from what you expected or how they're level of sophistication around this might be different than what you expected. Any surprises or any learnings from working with these partners? >> Yeah, so we went into these partnerships. I think a lot of people told us that like pharma doesn't know how to use AI. These are not. It's not like a tech forward industry and things like that. And to be honest, that hasn't really been our experience. I think that these are, again, they're very rigorous partners, very rigorous customers, right? And they're going to test all of our claims, right, before they start to deploy these things. And when they see the data, they go all in, right, because pharma is an innovation industry. And they tend to just think about even the whole economics of the pharma industry. If you build a product in pharma, like a drug, right, you only have exclusivity on that drug for a certain period of time, right? And you have to continue to reinvest. Eli Lilly is a trillion dollar pharma company right now. If they don't get more blockbuster drugs, they will not be a trillion dollar pharma company forever. And I think that forces these companies to really be on their game of adopting new technologies and deploying them and trying to stay ahead. Pharma is a very competitive arena. There's a lot of people trying to bring these drugs to patients with, by the way, it's great for patients. But it means that you have to be on top of your game here. And I think that means that, you know, once you're through that door and your models are working, we've seen like an upscale this adoption very quickly. And people think you might have to use the models in incredibly creative ways. I really like the incentive structure that it creates as well. Like taking more of the partnership model. As our models get better, we get better results to our customers and so on. So I think like that's just like a nice, like nice side effects for us as well. Chai has been like incredibly focused. And like really as the models get better, you can kind of like iterate on a better model with better data and so on. So like a lot of what we see and like a lot of what our partners are using the model for like generating, you know, binders, antibody, so on. Like we're also like doing a lot of dog food in house. Like we have a whole science scene that's using the models and just trying to understand what they can do. And a lot of that comes back to like, okay, now that like the model, now that we've unlocked to use case acts, like what type of new data can we generate? How can we make the model better that way? And really like that's been the long term vision of Chai as like we never thought of like we're going to build this one model that's going to solve all this. Like we know that they're just like in other fields, like they're going to be multiple iterations of this model. And as you get better and better models, like you get this flywheel effect, I guess. And you mentioned what data can we generate? Can we touch on data for a minute? You guys can't exactly just go scrape the internet and have all the data you need to build your models. Where does the data come from? How does it kind of compound over time? Can you just say a word about data as an input to your models? Yeah, so I think like the primary source of data or like the gold standard source of data is this like protein database. So this is like it's actually just like actually like legit lab scientists like who since like 1970, I've just been depositing crystal structures of like proteins and other molecules over time. And like really without that like structure prediction and design wouldn't have been a thing. So there are these. So that's like one source of structural data. When I started in the field, I really took like a structure pill approach. I was like super interested in predicting structure. Like how do you think about a machine learning model that can output 3D coordinates? Like that's that's not doesn't look like an LLM. It doesn't look like an image model. This is really in its own class. Josh interestingly was taking like the exact opposite approach. So he he was like one of the original authors on ESM and that was like a really like a seminal work and understanding how to apply language models to protein sequences. What's really cool about that work is if you can train a language model to understand protein sequences, what ends up happening is it ends up kind of like representing the 3D structure internally. And there's like a really interesting reason why this happens. So like in order to predict like missing amino acids, so like the same way that you'd predict like next word in a sentence, I want to predict next amino acid in a protein. In order to do that effectively, like you really need to understand like, okay, what does amino acids like immediate micro environment look like? Because that kind of tells you, okay, what are like the compatible amino acids with everything surrounding it? In order to do that well, you need to understand the proteins 3D shape. So I was taking like this this really structure base approach and Josh even more bitter lesson pilled the maze like we're just going to like look at the sequences and this is just going to emerge. So yeah, the two main sources of this data again, like protein database for structures and then like these massive, massive, maybe even like order of like trillions of token sequence databases. And what you can do once you have really good models is just run them on the sequence databases to get new structures out. So again, like you have this compounding effect as your models get better, they get more and more accurate at predicting these structures. And then you have more and more training data for the next series of models. Sometimes people ask us that shy about these days that shy, like which paradigm are you actually going after? I think one of the things I love about our team is that it's actually just like neither we're very pragmatic. Like we want to solve the problem. We don't really care. Is it a sequence approach is a structure approach and practice it's going to end up being like both of course. I think it'd be surprised, but there might even be more like biological sequence tokens on the internet than like English language tokens. And now a lot of that data is not that useful. It might not be redundant. It might be very noisy. But there's a lot of data out there as a lot of art and like bring it together. And I think also the exciting thing is as the models have gone into a point now where we can design things in the lab right to new targets. For instance, we can actually use the models to generate data as well. So there's a lot of exhaust from like all the experiments that we're doing at CHI, which also helped to make the models even better. So it's a similar kind of take off that we saw at LLM. Like when I was at OpenAI, we worked on reinforcement learning of like a GPT-1 architecture. Like did not work because the models weren't like good enough. But once the base model got good enough, then you could start to do those kinds of experiments. And I think there's a similar analogy that's starting to happen in our world now as well, where the models have reached a point where there's actually a renewed. Interest in data and like how do we actually bring the models into the loop on like making that happen. And then that creates another really like interesting cycle on compounding improvement of the models. Yeah, we're talking a bit about compounding improvement in the models and you said something about how pharma as an industry is an incredibly competitive landscape. And it's interesting because I think there's been a renowned interest in using AI and using ML to generate proteins and molecules. And so your arena has actually become quite competitive. And you guys have obviously done an amazing job at staying at the frontier. It's been the year of deployment for you. You've locked up a number of pharma partnerships that are making your models better. But how do you guys think about the competitive landscape and staying at the front of the frontier? Well, first of all, I think it goes back to like looking at the results and not fooling yourself and being rigorous. So one of the reasons why we do a lot of evaluation of the models. We mainly do it just to hellclimb in the models itself. I think if you look at many of the capabilities that we've brought and have kind of been like first in the field, you look at our try-to model, right? Like getting to success rate of the models that you didn't have to do large libraries, screening any more to see results. A couple months later, showing how we could bring in a lot of these developer properties. We've talked about before, like the manufacturer ability of the molecules. So in many cases, we are pushing the model forward and trying to see these capabilities emerge. And then we try to quickly like lock those in. Like how do we make those capabilities like even more pronounced that they become production ready and we can like ship them to our customers? I think one of the things like to tell the team is that, you know, it's not like we are head to head like with other model providers like that. Where all of us, I think, are working against nature. Like nature has actually been a pretty good baseline. People have done drug discovery a certain way for a long time. And you know Matt talked about how like maybe you don't want to add like module 24 to the Chi model. But people have added like module 240 to like the existing wet-lap protocols and they have been like tuned quite considerably. One of the scientists on our team Andy Young was, he was one of the first people working on yeast display at MIT actually like two decades ago. And he's got 20 years experience and like Pfizer and Genentech like really honing in these methods has a drug approval to his name and antibodies. And I think you look at someone like that and like, you know, and he knows how to make a good antibody with existing tools. And that is actually the bar that we need to clear. Now of course, I think the ceiling and AI is going to be a lot higher than what we've managed to do before. Otherwise what would be the point of doing this? We didn't start the company just to make you know a 10 times faster mouse, right? We started this to make breakthrough medicines that weren't possible before. But ultimately that is the bar that we needed to clear in order to get adoption. I think we hit that inflection point a couple months ago. That's why you've seen a lot of these big pharma announcements. But now we just need to continue to hone in on like making these things even better. For somebody who's listening who thinks, wow, this sounds pretty cool. I wonder what it would be like to work at Chi. What is the best thing about working at Chi and what is the worst thing about working at Chi? Yeah, I can speak to some of this. I'll thank you. on the fly the worst thing. (laughing) So I think the best thing is just how, actually, mission driven, everybody is, everyone is so dedicated to what they're doing. I've worked out other companies. The closest I've ever seen to this is maybe some of the guys in my PhD lab, but everyone is just incredibly, incredibly motivated. We all work really hard. There's an obvious shared goal. And I think that's really rare to see, and I think this goes back to just the focus that we've had since the beginning, and we've always had kind of a clear philosophy, a clear plan on how things are going to get there, how things are going to get better. And really, like, every night Chai is very bought into this. It's pretty amazing just to see the amount of dedication that everyone's putting in. Least favorite thing about Chai, not directly on top of Daniel and Chocolate, maybe. (laughing) - Maybe the next office. - Yeah. (laughing) Actually, so probably the least favorite part now is just like, I guess it's getting things to work at scale. And really, I didn't even know what that meant. Actually, when we started Chai, we had 128 GPUs, and I was like, this is the most scale possible for a lab. This is crazy. I came from my PhD group where they were four of us sharing eight GPUs, and I was like, I just felt GPU rich, it was crazy. So now, at Chai, we have a lot more infrastructure maintained. We have a lot more compute resources, like luckily GPUs are parallelizable, but that also kind of brings up its own set of problems. So just like, how do you keep a cluster healthy over time? How do you get that large training run? So how do you keep that running for months on end? And even when it does die, how do you automatically resume these things? How do you keep all of your communication down? How do you optimize the models and make the best use of the resources that you have? So I think these are a lot of problems that, you know, they continuously pop up. They're good problems to have, but I think they're also really difficult to solve. Yeah, and I'm just excited to work on this. Yeah, we, I remember one of the CEOs that we've worked on a couple times, Frank Slootman had this line about, you either have the pain of failure or the pain of growth. You'd much rather have the pain of growth. Yeah, that's exactly right. I think my favorite part is probably the results, and that sounds a bit cliche, but there's nothing. That's working. It's working, right? And just like knowing that you're, you know, I think many of us in the company, right? Like we've been working in AI for a long time, right? There's always experience you built up. And to know that you're applying to something that really matters, like just even, I think just take Matt and, and like we've been working on this problem for like 10 years, right? And a couple of years ago, I'm looking at some like, yeah, we were writing some cool papers, like everyone is like celebrating this. We're actually making the world better. Like is this actually going to impact some patients? And I think now the answer is like, actually, yes. Like we have reached the point where this is going to make a big difference in the world. And every time you get one of these big results, anytime there's a new feature on the product that makes our lives of our customers easier. Whenever there's like new lab results coming back from the science team, it's just always so, honestly, it's a gazilar rating to realize like, this is actually going to change the world in a pretty profound way. I think that's also then maybe comes to the least favorite side, right? Like, you know, we're running the company and it's like we have real partners that are relying on this and like things have to work, right? And you ship a new model generation. How do you make sure that there's no bugs in that? How do you make sure you don't have regressions, right? This is no longer just like a, again, the blue sky research problem of like, oh, we got some cool results to move on. We've had to have really high priorities on like, you know, having production level code bases as the team grows. How do we make sure that the code base is in a state that more people can contribute to this? So something that one of our other co-founders, Jack likes to say is that if you want to move fast in the long term, you sometimes have to just move a little bit slower in the short term, right? And make sure that you are building something, again, that goes back to that compounding idea, goes back to, we don't add module 25 to make the next thing work. So sometimes, you know, you're like so excited to get to the next result and you just want to jump into it. But we have real partners, some of the biggest companies in the world that are now relying on us. And it's important that we realize that we take that responsibility to heart and we make sure we're building systems that continue to work. - I have a burning question. Why are you guys called child discovery? - Chemistry and AI. But we love child tea as well. (laughs) It's a lot of child themed stuff in the office. - That's right. - That's a good question. I didn't know that either. I was all angry. (laughs) - It was a very user friendly name. - Yeah, we also wanted a name that like, biota companies have such complicated names. - Yeah. - We wanted something that's gonna be much simpler. We're trying to make this whole thing simpler, right? - Yeah. - So we need a simple name to go along with it. - Awesome. What do you guys most excited about in the next six to 12 months? - I think for me, it's just the deployments that are happening. So we've announced a couple of these partnerships and I'm really excited just to hear about the results that our partners are bringing online. It was back to this point of like making a real difference and also why I'm so happy to see how these partnerships are going even, you know, post into the agreement and as we're working with these folks, that the models are not just sitting on a shelf somewhere, like they're actually being used on real programs. People trying to approach devastating diseases where, you know, if Chad could give them a molecule that works, it could really change the lives of patients. So I'm really excited to see how that goes. Just the pace of progress here is incredible, but also just like the pace of the models. Like a year ago, you could not zero shot a molecule and like, you know, have a good sense that your program was gonna work. Like now that's changed. So I might zero shot a molecule and be like, I think we're gonna bring this program to the clinic now. And then a year later, you know, they might even have some of those first molecules going into patients. And just the speed of that is just incredible. And it's, you know, sometimes you get some shivers thinking about this that like, okay, my model is gonna, like Matt has some patents from our last company about, you know, like just generating the molecule on the computer and these things are now in patients and just to think about the scale, like, I don't know a few years from now, do we have dozens? Do we have hundreds of like, chimalculals going into people? It's a bit mind-blowing to think about what that might look like for patients. - I think one of the things that like, again, like what made Chienique like the dedication, like people are like, man, you work a lot. And like, you, like, aren't you burn out or whatever? It's like, it's actually really easy and like, it's very motivating when you're making progress. Seeing the progress they were making and just like thinking, man, the next model is gonna be even better than the last we identified this new thing, so on. Like that's so incredibly motivating. For me, it's more of just like, what can we unlock next? And like, how do we make these things more controllable? And like, when someone comes to us with a certain target, instead of just like, hoping we get a finileer or something like that, you know, can we actually control this? Can we say like, we want exactly a 10-nanimal or binder, things like that? Like, there are a lot of technical things that I think are, we're like, kind of right on the brink of solving. And for me, it's like, it's really motivating just like pin those things down and just like get all of this over the line and see kind of where that leads to next. - Awesome. Matt, Josh, thank you for engineering biology and thank you for sharing your story with us today. - Thank you guys. - Thank you for having us. (upbeat music) (upbeat music)

Podcast Summary

Key Points:

  1. Chai Discovery aims to make drug discovery more like engineering, using AI to design molecules rather than relying on trial-and-error.
  2. The field has evolved from protein folding (e.g., AlphaFold) to generative models like diffusion, enabling de novo molecule design.
  3. Key breakthroughs include deep learning for structure prediction (2018-2020) and diffusion models for generating proteins (2021-2024).
  4. Antibody design, once considered too hard due to data limits, is now a focus, with models achieving higher binding rates.
  5. The team blends AI researchers, engineers, and antibody specialists, prioritizing simplicity and high-quality software.
  6. The goal is to shift from "discovery" (screening millions) to "design" (specifying desired properties), potentially increasing lab testing due to higher ROI.
  7. Recent models show improved hit rates (from ~0.1% to higher), enabling more therapeutic-relevant molecules.

Summary:

Chai Discovery, co-founded by Josh and Matt, is a foundation model lab for biology, aiming to transform drug discovery into a more engineering-like process. The core idea is to replace serendipitous trial-and-error with AI-driven design, where models can generate molecules with desired properties, similar to how code generation works. Key milestones include the advent of deep learning for protein folding (2018-2020), which evolved into generative approaches like diffusion models, allowing simultaneous generation of protein structures and sequences.

This shift enabled tackling harder problems like antibody design, previously deemed infeasible due to data scarcity. The team, initially composed of AI researchers, has expanded to include engineers and antibody specialists as models improved. They emphasize simplicity in model architecture and high-quality software, with a small, over-capacity team to prioritize impactful work.

1% to more), moving toward zero-shot molecule design. This could increase lab testing ROI, paradoxically boosting demand for experimentation. Ultimately, Chai aims to industrialize drug design, making it more iterative and design-oriented, with the lab serving as verification rather than discovery.

FAQs

Simplicity is a core guiding principle. The team aims to simplify complex models, like Chai-1 with 23 submodules, to identify what's truly important, making research and scaling directions more manageable.

Chai Discovery aims to make drug discovery more like engineering by using AI to design molecules. They want to shift from trial-and-error to a design-oriented paradigm where models can materialize desired molecular states.

The boundary is shifting toward more design, but lab testing may increase due to higher ROI. The key is changing the paradigm to be more design-oriented, not necessarily reducing lab work.

Major breakthroughs include AlphaFold 2 in 2020 for protein folding, and later diffusion models enabling simultaneous generation of protein structure and sequence, allowing more realistic prompts with constraints.

They saw progress in antibody-antigen structure prediction, which was previously thought too hard due to limited data. This signaled that antibody design, a key therapeutic target, was becoming feasible.

Diffusion models work by iteratively improving a noisy, broken protein, teaching the model to make it slightly better step by step. This approach proved effective for biology compared to earlier methods like variational autoencoders.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.