Episode 16: Why Chemistry has Held Drug Discovery Back
37m 36s
Stan Yastoromsky, co-founder of Molecule.1, discusses how AI can revolutionize chemistry by addressing its data scarcity. He explains that unlike protein folding, which benefits from the Protein Data Bank (PDB), chemistry lacks comprehensive reaction databases, as scientists only publish successful experiments. Molecule.1 bridges this gap by operating a high-throughput laboratory that generates massive, high-quality datasets (100x larger than typical) for specific reaction subsets. This approach enables AI models to learn superhuman correlations between molecular structures and reactivity, overcoming chemistry's inherent complexity. The company's platform, Maria, integrates reaction execution, model training, and experiment planning to discover new chemistry and accelerate drug development. Stan highlights that scaling laws—where performance improves with data size—are crucial for chemistry, and that cultural resistance (not technical difficulty) has hindered prior efforts. Data quality is paramount: Molecule.1 built custom ML software to analyze diverse reactions accurately. By combining dense, reliable data with AI, the company aims to reduce drug development timelines from years to months, making chemistry as predictable as printing. This vertical integration of lab automation and deep learning positions Molecule.1 at the forefront of autonomous discovery, with potential to transform medicine and materials science.
Welcome to From Models to Medicine, a podcast from Cammy Think Tank. Exploring how AI is actually being applied across the life sciences industry. Wear your hosts, Kamayani Gipta and Michelle E. And this is From Models to Medicine. Hi everyone, welcome back to another episode of From Models to Medicine. Today I am joined by the co-founder of Molecule.1. I am here with Stan Yastoromsky. So Stan, thank you so much for joining today. Really, really grateful that you're here with us. And you know, before we even jump into the questions, give us a little bit of background on who you are, what you do, and why you're here today. Thank you for inviting me excited, excited to be here. Maybe a few words of as a form of introduction. My background is in deep learning. And kind of an understand that most of the audience is in life sciences. So I think maybe a bit rare though now, maybe not as rare, case of someone immigrating kind of from computer science to, let's say, life sciences right in my case. It's mostly synthetic chemistry. So yeah, excited to be here. I love it. I love it. So as you mentioned, you you started in deep learning research and then you've ended up co-founding a company at the intersection of AI and high throughput chemistry. So is there a particular moment or a particular problem that made this pivot inevitable for you to pursue? Yeah, I mean, that was like a pretty specific moment in my life. So kind of the starting point is me going into deep learning. So I think I started into deep learning in pretty early days of the field. So I was always drawn to weird ideas and at the time that was deep learning. So I was fortunate enough to work with the founders of the fields. So you should have been doing particular. But then I was a kind of increasingly drawn into life sciences for several reasons. My last step before molecule one was postdoc at NYU. They did postdoc which focused on breast cancer and AI for that and also optimization. So I always was very much interested in understanding how deep nets discover all of these capabilities. And then it was like very interesting time for me in my life because the pandemic hit in 2020, right? So New York City was definitely very weird at the time. And also in the background of this are like kind of in that moment also the opportunity to become so called late co founder of molecule one and meshed. So the other co founder reached out with the proposition for me to join as executive team. And I got really intrigued. So what molecule one at the time had as the core idea. Again, thing that I would say is was very rarely done was to do a for chemistry, which in terms of deep learning specifically for chemistry that didn't exist as much. And core idea was to focus on the data, even more so than models. So in. Probably no better than me that in protein folding what was absolutely essential was so called protein data, right. PDB and nothing like that exists in chemistry. So in chem, when we train those neural networks in chemistry, we usually go off chemical reactions recorded in literature. And I think it's most audience knows we scientists generally publish only positive outcomes, not published failure. So it's really difficult to train your networks of a bunch of successful experiments without really the whole breath of what doesn't work, which is the critical part of knowledge. The unique proposition of molecule one, which was really intriguing was, well, how about we set up a hype throughput laboratory that's fast enough to build towards something like PDB, but for chemistry and then trained deep network. So, you know, deep networks train on big data discovering new things about chemistry that was really interesting to. Wow. Yeah, you're absolutely correct. I mean, you know, for protein folding, you know, looking at the database of that, that is something that is very clear to biologists who are in the field that this is like the area that we're going to go into. And from what I'm hearing from you is like in chemistry, that didn't really exist. And so figuring out what that could look like that becomes really important as you are doing model development, because if you don't have the correct, you know, the right sets of data formulating, then it's going to be really hard to train your model, if I'm understanding correctly. Yes, I think it's even goes a bit beyond that. And in a way that PDB in some ways, doesn't so our, the make like we name our platform that brands the reaction, trains the models, plans experiments, Maria. After Maria Skodowska-Curie, she's Polish, French, double loaded in chemistry, of course, right. But why did we name it like that? It's mostly because we really wanted to signal to mostly us that our goal is not only to automate chemistry and build these data sets, but go beyond what is known and discover new chemistry. And I think setting up a company that is vertically integrated in the sense that we both train your networks to serve our clients. Yeah. But also discover new chemistry. I think that's the interesting thing because we might be able to go beyond what is in the literature. Not only filling those holes of negative data, but discover new reaction files, discover a new set of conditions and hopefully eventually make faster medicines because that's kind of what is the goal ultimately. I love it. I love it. And I think you a little bit alluded to this is that, you know, chemistry doesn't get the same type of attention, right, as biology or genomics that it is even though chemistry underlies every single facet of our existence, right. Everything we got is an agriculture material sciences. What do you wish that people understood more about chemistry in terms of like why chemistry is the bottleneck that actually matters within these processes? Yeah, I mean, you said it perfectly. It's amazing how much runs and chemistry from things we rely on from materials, like we culture, it's just something that I don't know chemistry doesn't get good PR. So I think core idea to have in mind is that there is a huge gap between what you want to make. So you want to make a medicine. You imagine a perfect molecule that is designed by your great, generative models, but there's a huge gap between that idea and what you can make in the laboratory. So chemistry is just really hard. Yeah, quantum physics, you know, temperatures, mixing, like all of that. And it's really hard for humans to predict outcomes of those experiments. So at each stage of the process, we compromise between what we would like to have and what we can make. And when we maybe don't compromise because we really try hard to make something, then it really prolongs the time. So I think making chemistry a bit more towards like printing in the sense that we just print what you want. It's not like you have a document and you're like, I want to print it. It's just to complex. Unless you want to save on ink, but generally no. I think then we can go from like years, which now it takes for molecule to going to clinic to, you know, it can be like months because ideas are pretty good. Like you're getting pretty good with models, at least of some biology. But chemistry is a bit of a lager. I would say. One of the things I found most striking about the work that you guys do at molecule is, you know, because we know that possible molecules vastly outnumber everything ever synthesize and quantum physics makes predicting that behavior really, really hard. How do you think like AI can start to make a dent in a problem that is so enormous that even, you know, for us within like the biology sector, like it's a very daunting problem to take a look at. And so how can you start leveraging AI within the space? Yeah, indeed chemistry is really hard. So there are two components to that. One is discovering new reaction classes is actually the least less obvious, but I think potentially more impactful where the idea is maybe we can discover chemistry that is just more robust, better. Maybe it's matter of, you know, if you cannot play a game, change rules of the game. And one Nobel Prize in chemistry went to something called click reaction, which basically it's called click reaction because it just always works. So that's one area that I am pretty enthusiastic about. And in, by the way, biology is created that because it designs these enzymes that make reactions extremely efficient for specific substrates. But well, most of our work is with existing reaction classes where we discover some kind of reliable processes.
So here we had this like a ham moment because when we open our laboratory, it was in 2022. So four years ago, roughly, it was a huge bad at a time. And they're like, pretty scared, I would say, because we figured that, okay, if we're really believing that in the story that we need this tail data, then it means we need to start generating data sets, 100x at least larger than the largest reported HTE, some high-frippled experimentation data sets of chemical reactions. So at the time we honestly didn't know, and what I will say now is, one of the empirical proof that we discovered that this works. So where it works and where we can tame this complexity of different molecules, is when we fix some of the variables. So the thing is that in our laboratory, we perform experiments in a very repeatable way. So we focus on sub-sets of chemistry, and for this subset of chemistry, we run a lot of reactions. And that allows the model to learn from just dense data, correlations between structures of the molecules, and reactivity that are superhuman. So superhuman, in the sense that they escape human understanding of chemistry. And we compared to humans, our models, and they are much better when designing reactions specifically for the Maria platform. But that's the second thing, right? Like, again, the kind of change in rules of the game, it chemistry is so big, yeah, the quantum physics is not yet possible to calculate well. Then focus on a subset of chemistry that's extremely important, and generate a lot of data for that, and kind of try to make your process of generating data very economical and fast. Yeah, no, that train of thinking, like the chain of thinking in terms of, yeah, focusing on that subset, right? Because it's such a large data set, that makes sense because if you narrow the scope of what you're exploring, then it becomes a lot more reasonable to dig into what's going to go into it. And so, as you mentioned, molecule one is like one of the, among the first to generate this large scale data specifically from chemical reactions. What did that decision to go after that real laboratory data unlock? Like, what are some of the discoveries that have found? And why haven't other people done it at scale before? What's the reason? Yeah, so maybe first on like why others haven't done it. So the company started more formally in like 2019, so a bit later when I joined us, late co-founder, and back then I was an advisor. And at that time, we were working on declaring it for chemistry. And basically, no one was commercially, I would say, to some simplification. Like we built the first commercial deep learning software for synthesis planning. So the task of strategically figuring out how to make a molecule in the laboratory. And only from the deep learning perspective, opening such a lab made sense at the time. Because it wasn't part of the typical training of chemists, runs so many reactions. One of our chemists told us that she runs weekly, more reactions than normally a chemist to run like a whole career. And it's just against the grain of training of chemists. But very natural to a company that was pedigrees in deep learning. So I would say it was mostly cultural, just interesting. Often companies live for a day by its culture. Like, you know, you can, I can give you a very good example. Maybe a good example. Yeah. So Apple, I don't know, I probably noticed that Siri from Apple is not like super intelligent. It's interesting because now there's this deep learning and stuff. So why is a good question? And it's a culture thing. So as far as I know, Apple really, really cares about robustness and just reliability. So they made a decision not to go into generative AI for Siri, at least for the time being. So these decisions can be extremely important. And we made a decision that it really matters to build those large data sets. So those really cultural, I would say. And now, as the world has turned for into deep learning and leaned into that philosophy, their company is that also generate a large data sets of that kind. But I would say we're the first to leaning to that very heavy. As far as I know, of course. Yeah, of course. And I think even having the vision back in 2019, you know, you joining the team soon after, I'm bringing that deep learning capability to examine those large skill data sets. I mean, that showcases the ethos of kind of like the visualization of what you want to build on it. I guess that takes me to my next question of, you know, high throughput experimentation has been a mainstay in biology for decades. Like it is, we all know about it for those in the field. But from my understanding, it's a lot harder to pull off for chemistry. And what I'd like to learn is like, what has changed technically? That makes high throughput chemistry at the like the microleaders skill possible today. And where exactly does AI fit into making that type of experimentation work? Yeah. Is it a very interesting question? So I wouldn't say necessarily that chemistry high throughput is like fundamentally much more difficult than bio-watching. Maybe in the sense of physical operations. So I think large portion of that was cultural. But there are elements to that which are unique to chemistry. So in biology, that's pretty technical but very often we work with, you know, liquids, let's say, of relatively low concentration liquids that you move around. You have pretty, let's say, empirical tests standardized analytics. It's a constrained workflow in some sense. Of course, over-centrifying grossly. But in chemistry, you more often have more difficult physical environments. So you have high temperatures, you know, slurries. Sometimes you have reactions where you need to like shine light on them to happen. So there are those difficulties. But in our case, we focus very much on chemistry that it looks like biology in the sense of like these homogeneous liquids. So the making of, let's say, reactions or creating the data set itself, that part kind of was more of the cultural side that I mentioned. However, what was difficult is that you really needed to scale this data set size to start to build models that understand the chemistry. So that part was less of it. So we in declining there is this thing called scaling loss. So how your performance increases with data set size. And we made the kind of the bet that scaling loss exists for chemistry. And it really matters to go to the large data set size. And few years after setting the lap now, as we can plot this scaling loss, it's really true. So while setting up these reactions is kind of doable, it's difficult but kind of not very much the more difficult than in biology. The difficulty is then predicting outcomes of reactions done in such high throughput manner. And it is difficult for humans as well, right? Like in particular for humans. So a lot of data is needed to predict those outcomes. One of the reasons for that is that to do high throughput experimentation, you often need to go to lower volumes and sometimes change a bit of chemistry in ways that are less obvious to chemists. For example, you might want to use specific solvents, you might want to use lower temperatures. So there are things in which you have to adapt your chemistry, which make it harder to predict for humans. So I would say like to drop up the answer because there's a bit technical, that none of this thing in chemistry was the really the need to go to big data sets. I would say to really land those models. And maybe in biology, maybe it wasn't as necessary or maybe there are different problems I don't know. But in chemistry, you have to go big and no one really tried as much as it was technically possible. Sure, sure, sure. And I think it's interesting because there is this debate between like, do we have large data sets? Can we train models on smaller data sets? Which direction should we go into? It's a huge debate occurring in biology right now. And is that debate occurring in chemistry as well or is it that you just have to have a large data set in order to have models work appropriately? Yeah, that's a good question. Well, my take is that if you can build good quality big data sets, and that's not a huge compromise, I think it's usually worth it because there are things that you get from scaling laws and understanding that I have to do with small data sets and just to learn it shows us that scaling laws are really important. That long term, I think, strategy that some companies are kind of exploring is trying to think about this as not binary. So you can have some data that is lower quality but bigger, some that is smaller and higher quality. So, for example, now as we do more projects towards autonomous discovery, we really lean into our ability to build larger data sets. But then we can dig into the pockets of like, "Oh, I found something interesting." Right now, we kind of see that that when we go big, we better often discover something unexpected and then we can
kind of start to do this like, let's say less of a white chemistry type of thing, but like digging into like all kinds of understand that that small experiment that was unexpected, right? Yeah, yeah, yeah. No, that makes sense. And I think the quality of the data is going to be the most important thing, you know, beyond I think that's one thing that we're finding within the life sciences area. I was having a conversation with someone who looks at rare diseases, right? And they have a very small subset of data that they're able to produce, but that data is very high quality. And the models that they're building on it is able to create a loop mechanism in which it's taking the high quality data and it's able to create synthetic data that is of the same quality in order to assess like, okay, for this rare disease where we won't have X amount of more patients because it only, you know, is impacting a small subset of individuals like this is additional data that will help us develop kind of like the mechanisms of what is being shown for those diseases today. And so it is interesting always I that is like the number one thing is like great if you have large sets of data, but the quality is going to be really, really important as we train these models within the problem. Yeah, definitely. I mean, what took us quite a long time is figuring out the quality part. So one thing was that we had a lot of reactions, so this is like terabytes of data, overall data. And there was no software on the market that would analyze it accurately. Not because of the size, but because of the diversity of the reaction, it's not of a technical thing, but there was no software. So we had to build our own software. It's probably the first software for analyzing this data that uses machine learning insights to analyze the data. And all of that was those years were spent really on increasing the quality of the data. So so that is absolutely critical because in the end, we want to produce the molecule that someone wants because that might be the medicine. So we really care about this precision. Yeah, no, that's helpful to understand. I know you talked a little bit about autonomous, you know, the frontier that's opening up right around autonomous research. We are seeing that LLMs are now being combined with automated laboratories, which is so cool. It's very sci-fi, you know, something I did in vision 20 years ago. And we know that this combination is generating and testing scientific hypothesis without even having a human in the loop for every single step. Where is that actually working in chemistry? And where is it, I guess like, where is it hype versus reality? Like what is the actual use cases that you're seeing in chemistry today? Sure. So I can unveil a bit. I cannot say unfortunately everything, but we have started recently with a frontier where we love very interesting projects exactly in this line. And I would say there are definitely limitations and we see it is actually very interesting. Like through this project, we understand we are in the frontier of capabilities of LLMs. We use, you know, as good LLMs as you can get. And like we see very precisely what's the slacking. But no, I mean it's, I would say it's more on the side of reality than hype. So let me explain, you know, you know, actually, I mean most people who know me would say that, I am the skeptical guy, but maybe I'm biased because I'm from deep learning work, but I am very excited about this. So what we are doing is leaning towards the fact that we can run a lot of reactions. And then with the agent that we compute, we generate hypothesis, and then we test them in our laboratory where the hypothesis are about, well, potentially in your chemistry, right, like publishable in your chemistry within one month since signing the project. We completed the whole one cycle of like, you know, designing the experiment, running it. And then most interestingly, getting the results, feeding it back into the agent. And I think that how moment was for me when it was reasoning about the data in way that was like, you know, solid. It was like really interesting. I cannot go on for something to like technical details, but it's not yet, you know, it's not yet autonomous or agentic fully, but what we see definitely is that it remove like opens up research to areas that human humans might be overlooking. And also it really speeds up research. So you ask earlier why chemistry might accelerate drag discovery. I think it's a bit like that. So using the lens in the scientific process might extremely speed up the process because we humans spend too much time deliberating some of the things, well, it would be better to run reasonable things but much faster and go towards more of exploration. So I am not enthusiastic about it, but as I was saying, we see limitations. It's actually what, like the reason I went into deep learning, this was that I was really fascinated by deep networks discovering those things, those capabilities. They learn first words. They learn how to, you know, glue them together. These are the things I worked with as a researcher in the field. So it's really interesting to see that the front here is that deep networks yet are not at the point where they can propose extremely high novelty creative idea. That's really hard for them. They also have very big limitations in scientific reasoning. And those limitations will be lifted, I think, but it's really interesting to see them crystallize in that application because they do not show up as much in other fields. Because in science, you need to work with new models of reality that might be underrepresented in your internet as the point of science. And it really drives these models to the edge of their capabilities. So it's really interesting. But we'll find the side of the promise how much it can speed up science. And I think we see outcomes that are scientifically interesting, but also on the side of driving all of us, like as a tool of deep learning towards age, like true age AI, because that's really the frontier of capabilities of these models. It's not yet, I think the words, the hype, is telling you that this is like fully autonomous. And, you know, like I think those limitations mean that we are not, that I think we need to go over these limitations to really have, you know, an AI, land, another price. I think those limitations have to be lifted. It's not just my, I don't think it's just matter of taking existing models. Yeah, I completely agree with that, right? Because if those limitations aren't lifted, then how can you get towards developing a complete like AGI that is able to fully be functional and work autonomously and not really need any like zero input whatsoever from a human or even like new research that's coming through. So that's helpful in understanding like how it's working within the chemistry field. And really excited to see what you guys end up doing, especially as, you know, you end up partnering with this like Venture Air Lab where you can do a lot more of this work in deep learning. They're cool. Yeah. Stan, a question for you. So, you know, we know that biologists can now design like near perfect molecules using tools like Alpha Fold, but then they have to like compromise because synthesis is too hard, it's too slow. Yeah. Things happening as AI models become better than humans at molecular prediction. How does that change the relationship between what's biologically desirable and what's chemically possible? Yeah. Yeah. So they're actually fast, fast through that. So, but definitely as you say, models have gotten dramatically better at understanding, well, mostly those things like binding, I wouldn't say they understand like pathways of activity and things like that. You know, the fact you know how to bind into a dopamine receptor doesn't mean you understand the cause of depression. You don't. Of course, right? You know that much better than I could, but yes, they are getting much, much better. And then chemistry because of that becomes the bottleneck and lifting that bottleneck might exponentially accelerate drag discovery. Chemistry is really important in terms of both the time and the cost. So, by some calculations chemistry to calculate also synthesis as part of the clinical trials and things like that can be easily like 30% of just, you know, capitalized cost of drag discovery. It's much bigger than most would would think, but more importantly time is most important. And that's that's really as we get to better biology, the prediction that chemistry starts to be more of a bottleneck. One thing I would maybe throw at because this I may maybe important for for listeners is that there is big limitation now on the models of even that biology like binding. But those models and it is a bit related to the limitations of LLM's for scientific discovery mentioned. Those models, well they will work very well and this is shown repeatedly in papers, including most recent papers about those type of models for binding activity of all being. For proteins and, you know, like pockets into proteins.
that are well represented. And as you move further and further to targets that are less, common, new targets, or as you move to different type of pocket, not just the main one, than all of the sudden the performance drops, but well, it's obvious it drops, what is less obvious drops catastrophically quickly. So it's really interesting because if you think about this, everyone now on this planet is doing what I would say is maybe a bit vibe, vibe binding. So basically taking, you know, I don't know, bolts or something and producing heats or like, like, the design of molecules with those type of tools. But the problem is that then everyone is drawn to the same type of molecules because everyone uses the same binding tools. And I think that's a problem. I think we are losing creativity. And again, here chemistry can help because well, you can just explore more maybe instead of relying on the most confident prediction. So there's a danger in relying to match on biology models, just like there's danger in relying to match in GPD, such as the thinking. It's kind of similar. So just a word of caution about those models. Yeah. And I like that you tied it back to the creativity aspect of like experimentation and what people are looking for to your point of, you know, if we are looking at a limited subset of binding agents, you know, interacting with the work that we're doing, well, that's only going to make results that are perhaps like in the eyes of the scientific community. It is desirable. But then you don't get as many like negative use cases, which honestly, negative data helps so much in the development of the work that we're doing. You don't get so many outliers, which again, also helps with developing new experimentation and creating new hypotheses in other areas. We wouldn't have a lot of the modern day inventions that we have today had we just directed ourselves in one direction of how we're supposed to create these things. So I do appreciate that word of caution because I do think as we're finding, you know, in the creative field, for instance, like writing is feeling a little bit more boring or like you can tell us now. That's true. That's right. That's writing it. And I think we will reach a case in which even experimentation is going to feel very lopsided, studying particular components because we have valuable data, quality data. Maybe we have models that are built to scope for that particular problem. But we won't be able to explore those outliers. We won't be able to look at those negative use cases. And that kind of takes the fun and the creativity at a scientific discovery. And so I really appreciate you bringing that up as a word of caution to folks. Because it's important. I always love asking folks on this podcast, whether there is a person, a paper, or an idea that you've encountered recently within your field that has shifted how you think about where this field is heading in the future. I love this changing. You know, Ken Mr. Brooks. We are at molecule one in scientific discovery field. So maybe one idea or like being there that everyone can, let's say, play with. So what we are running out of in AI is good benchmarks. And recently there was launch of a benchmark that I really like and is really fun. It's called ArcAGI 3. So you can go to the website. It is interesting in that those are games. But those are games that are new always for the model. So maybe you maybe listeners know about things like AlphaGo. So playing Go using AI Alpha Star which played Starcraft. Those are always the cases where you play one game a lot of time and you're beat human by the sheer amount of experience. ArcAGI 3 is interesting because here you get a new game in tests that you haven't encountered. And well, why do I mention that will be in the context of scientific discovery. The reason for that is that there is a very deep connection. There is a benchmark that humans have dramatically better when they go to a new situation and they just like figure out rules as you go. You see that a lot in babies. babies have extremely interesting abilities to reason from partial data. So there are these fascinating experiments where they show how the baby effectively influenced like biogen statistics in like very early on like you know before it talks. It is reasoning out how surprising this experiment by done by a human but I was very maybe listeners for details. So ArcAGI 3 is simulating a bit that so it's trying to I think make benchmark for small scientific discoveries because what you need to at least in the early part of playing the game is discover how it works without any kind of bias. And it's interesting how AI is reported. I think that will be solved pretty soon but I think ArcAGI 4 I would guess would go even further inside to be discovered. That's my hypothesis. And I think this will be the frontier for a bit. So the idea I wanted to share with listeners is that it is really interesting time in AI. Where scientific discovery has become both the goal in itself but also potentially one of the final frontiers for AI. So I think it's really interesting times. And really that's the task where the critics of AI crystallize very well and you can see there's a lot of truth to them because this creativity you know kind of maybe to some extent last frontier of human intellect is really interesting. Some kind of really watching close to that. We at Moleculean try to contribute our part and kind of adventurous but we'll be as a field in a few years of time. No I absolutely love it and I would I'll find the link and I'll put it in the show notes because I'm sure other people would appreciate you know being able to experiment. Yeah thank you so much for sharing that and Stan honestly thank you so much for joining us on this podcast. It's been such a pleasure getting a different perspective hearing about the chemistry and of how it is you know bringing into the foray of drug discovery and what we're going to be utilizing for. I really appreciate learning about the differences you know between biological reactions, chemical reactions and how like they will be able will be able to bring these you know fields together in order to solve for better medicines for patients at the end of the day so I really appreciate you taking the time to join us today thank you so much. Thank you my cousin. Thank you so much for joining from models to medicine. Camille Think Tank is an educational organization that helps life science professionals learn how to use AI practically in their day to day work. We can't wait to see you next week. Subscribe to our podcast on either Apple Podcasts or Spotify today.
Podcast Summary
Key Points:
Stan Yastoromsky, co-founder of Molecule.1, transitioned from deep learning to synthetic chemistry, driven by a pandemic-era opportunity to apply AI to chemistry.
Molecule.1 focuses on generating large-scale, high-quality experimental data (like PDB for proteins) to train AI models, addressing chemistry's lack of comprehensive reaction databases.
The company's platform, Maria, aims to automate chemistry and discover new reactions, moving beyond literature biases toward positive results.
AI helps tame chemistry's complexity by focusing on specific reaction subsets with dense data, achieving superhuman predictive accuracy.
High-throughput chemistry is culturally challenging but technically feasible, requiring scaling laws to improve model performance with big data.
Data quality is critical; Molecule.1 built custom ML software to analyze diverse reaction data accurately.
Summary:
1, discusses how AI can revolutionize chemistry by addressing its data scarcity. He explains that unlike protein folding, which benefits from the Protein Data Bank (PDB), chemistry lacks comprehensive reaction databases, as scientists only publish successful experiments. 1 bridges this gap by operating a high-throughput laboratory that generates massive, high-quality datasets (100x larger than typical) for specific reaction subsets.
This approach enables AI models to learn superhuman correlations between molecular structures and reactivity, overcoming chemistry's inherent complexity. The company's platform, Maria, integrates reaction execution, model training, and experiment planning to discover new chemistry and accelerate drug development. Stan highlights that scaling laws—where performance improves with data size—are crucial for chemistry, and that cultural resistance (not technical difficulty) has hindered prior efforts.
1 built custom ML software to analyze diverse reactions accurately. By combining dense, reliable data with AI, the company aims to reduce drug development timelines from years to months, making chemistry as predictable as printing. 1 at the forefront of autonomous discovery, with potential to transform medicine and materials science.
FAQs
Molecule.1 focuses on generating large-scale, high-quality data from high-throughput chemistry experiments to train deep learning models, aiming to discover new chemistry and make faster medicines.
Chemistry is a bottleneck because there's a huge gap between designing an ideal molecule and actually synthesizing it in the lab, due to the complexity of quantum physics and reaction unpredictability, which prolongs the time to clinic.
Unlike traditional chemistry that relies on literature data with mostly positive outcomes, Molecule.1 generates its own dense, repeatable experimental data sets 100x larger than typical, allowing AI to learn superhuman correlations between molecular structure and reactivity.
The platform is named after Marie Skłodowska-Curie to signal the goal of going beyond automation to discover new chemistry, not just generate data.
While physical operations like high temperatures and slurries pose challenges, the main difficulty is cultural—chemists aren't trained to run many reactions—and the need for very large data sets to predict outcomes accurately.
Data quality is critical; Molecule.1 built custom machine learning software to analyze terabytes of diverse reaction data accurately, ensuring precision for producing target molecules like medicines.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.