Go back

Applying Google's Search Recipe to Pharma R&D Data - With Douglas Selinger from Plex Research

50m 55s

Applying Google's Search Recipe to Pharma R&D Data - With Douglas Selinger from Plex Research

In this podcast episode, Douglas Selinger, founder of Plex Research, discusses how AI is transforming drug discovery by leveraging vast, often underutilized datasets. He explains that his platform, inspired by Google’s search algorithms, integrates diverse data types—such as gene expression profiles, proteomics, and metabolomics—to reveal hidden biological connections. Selinger highlights a key frustration: scientists are often blamed for generating messy or poorly annotated data, but he argues that algorithms should be robust enough to handle data as it exists. This approach, he says, unlocks the potential of historical data that might otherwise be discarded. Plex Research’s platform emphasizes transparency, allowing scientists to access raw data behind AI insights, ensuring they are both explainable and actionable. Selinger also addresses skepticism about AI in drug discovery, noting that past failures have made scientists wary. He stresses that successful AI tools must be reliable and scientifically rigorous, offering immediate value to earn trust. By repurposing algorithmic principles from internet search, Plex Research aims to solve the long-standing problem of data silos and underutilization, ultimately accelerating drug development in areas like target identification, biomarker discovery, and precision oncology. The episode underscores the need for algorithms that adapt to real-world scientific practices, rather than demanding changes from scientists.

Transcription

9680 Words, 52885 Characters

English
Welcome to Tech & Drugs, the podcast where we explore how data and AI are transforming farming and biotech, Amu Hose, Tibogia Reed, scientists, technologies, and lifelong connector of science and data. In each episode, I see that we have innovators pushing the boundaries of drug R&D to uncover the real impact of AI beyond the birth world. Today we are excited to welcome fascinating innovators in AI driven drug discovery, someone who has been at the intersection of computational biology and farming innovation for decades. Douglas Selinger is a founder of NCU of Plex Research, a company pioneering a novel AI driven analytical platform designed specifically for scientists. What sets Plex apart is its commitment to transparency. Scientists have direct access to vendor line data, ensuring that AI driven inside are both explainable and actionable. Doug is no stranger to large-scale data analysis. As an early pioneer of microarray technology, he offered some of the first publication on experimental and computational approaches for large-scale transcriptional analysis. After earning his PhD in George Shirt lab at Hava, he spent 14 years at the Novotid Institute for Biomedical Research, Niber, working across the entire drug discovery pipeline from target identification and validation to high-fruput screening and pre-clinical safety. In 2017, Doug founded Plex Research to develop a new form of AI inspired by search engines algorithms. Plex platform integrates vats and disparate datasets, small molecule biodeactivities, gene expression profiles, genomic, proteomics, metabolomics, etc. To reveal hidden biological connections. By making sense of massive chemical biology and omics data Plex help biotech and farmer companies accelerate drug discovery in areas such as target ID and validation, disease identification, biomarker, discovery, safety assessment, precision oncology. So how is AI truly shaping the future of drug discovery and how do we move past the hype to deliver real, interpretable and actionable insight? Let's dive in, Doug. Welcome to the show. Thanks, Tiva. Thanks for the invitation. It's really great to come here and have a chance to talk to you about this topic. We'll talk about the company. We'll talk about AI data, of course, but I like to start with getting a little bit more history and insight on the person I'm talking to and understand how did you get to where you are today? What brought you to this field? Can you tell us a little bit about the early days? Sure, yeah. So I'm a block of the biodech is my training. So as you mentioned, I did my PhD in Harvard and German churches lab, which is a really fun place. There's a lot of cool things that come out of there. It's really exploratory, all kinds of crazy stuff. And I was there at a really interesting time when genome mix was coming out. So we knew genomes in the catalog of the parts list, all the pieces of the cell, all the genes, but the question was how does it all work together? What's how does it function? What are the dynamics? What's the can we do that on a large scale? So I was working with some of the early technology, micro technology, so looking at gene expression on a large scale, and I was in the lab actually getting that to work. And once we had the data, now we needed computation, we needed software, and there was no software at the time to analyze that kind of data. So I shifted to computational approaches. When I graduated, I went to the artist, when they first moved up into Cambridge, Massachusetts, and worked on all kinds of large scale projects, both as a computational biologist, also running a large scale transcriptional profiling platform. So both computation and experiment over many kinds of drug discovery projects. So I really got a flavor for all the kinds of approaches that are being used in drug discovery and the interface of large scale biology and computation in the drug discovery process. At some point, I had on this idea that we have all this data. So at a big company like the artist, certainly there's a lot of data. And in the public domain, also there's a tremendous amount of data just in the public space in the databases and in all these supplemental files that accompany all this scientific literature. And I hit on this idea that you had Google had tremendous success with the internet. So taking the massive amount of information on the internet and looking at using the structure of the internet, what are people linking to, for example, and using the things that people linked to to decide what's relevant, what's most important. Follow the links and see where they converge. And what I realized is that you could take that kind of approach and apply it instead of web data is apply it to data and ask across all the world's data, where does it point to a specific answer to a scientific question. And so is there more in the data than we're currently seeing? So does the data actually point us to answers that we weren't aware of that might not even be in the text of paper. So a long way to say there's this really interesting idea this new approach ended up leaving no vartis in 2017 to found flex. And here we've been building a system to really take that to use that on a tremendous scale. And we've been working with lots of companies at this point over 40 companies, Bataq and Pharma, helping with all kinds of scientific problems using this very general approach. That's if we look at the world of data, all the world's data and principle could be put into this sort of system, we can ask here's my question, what does all that data have to say about it? Right. So what's the best answer or the answer for the most support? And then what's the actual data that supports that answer or that discover? So we'll come back to Plex in a moment. But one thing I've seen some time with entrepreneur is that they start something out of frustration in a role that they had. So I used to work with someone who was managing data at the large pharma and the data was a mess and then he was like, okay, we need tools to fix that. And if we have this problem here, most probably everyone has the same problem. So was there any specific frustration that drove you to try to fix this frustration? Oh, that is a great question. Yes, tremendous amounts of frustration. So there's a couple categories. So one is, so I've seen so many panels and I've been on some of these panels too, but mostly seeing these panels where people talk about data silos that we have all this data. And we want to somehow use it to do computation and to make it machine readable and usable on scale and all of these things. And the panel comes up with, if only the scientists, because they put the onus on the scientists, if only the scientists would generate more data, clean it, or better data, the right type of data, the better annotated data. It's always the same, it's been the same list of complaints forever. And it's not the same we haven't made any progress. There's some that we've made some progress, but it's still the same issues. And what the point that I raised is that maybe it is the algorithm is maybe we need some algorithm that can handle the data as it is, right? The data we have now. You know, I've never gotten a call from Google saying, your website, could you take down your website? It's really screwing up our algorithm, right? No, it's it's robust to noise into whatever people put on the internet doesn't it does not impact or has very little impact on the results. So you need an algorithm that works with the or an approach just generally that works with the data that we have. You can't currently wish that we had different kind of data. And if you let me the other big source of frustration is that we have this disconnect or we've had for a while now a disconnect between our publication process and our technology. So for a long time now, we have we generate huge amounts of data. And I think scientists know this. I'm not sure how aware the rest of the world is. There's certainly scientists in this field realize this is happening. Perhaps we've compartmentalized it and stopped paying attention to it, but we write papers, we give presentations and there'll be a slide, let's say, with thousands of data points, a scatter plot, a heat map, we have these dimension reduction plots, you map testing all these crazy things that have thousands and data points that sometimes behind them each data point represents a thousand data points. So we routinely just show these massive data sets, but the amount of duration we extract from that data is tiny compared to the amount of data that's in the figure. So there'll be a figure with a thousand or a million data points and then we'll summarize it with a set insert to. So somewhere we know that we're missing out that we are not capturing even the tiniest fraction of what we've collected, the data we've collected, the tiniest fraction that is actually turning into knowledge. And we know that, but I don't think as a field, we know what to do with that. And that was another big source of frustration that I was sitting through these and like, and I would see this kind of presentation and realize that wow, the data set that was generated right there contains so much information that we're not leveraging. This is a huge opportunity that just being wasted because we just don't know what to do with it. And both of those those problems that I laid out are things ultimately that plexus is dressing. And I think we really haven't approached that fits both of those problems. And so it was truly worn out of frustration certainly. But I worked on a few projects where we were looking into a historical data set to train models. So we had collaboration where we had companies that had very interesting predictive model, but of course to improve those models, you need to friend them with a lot of data. So we were going into our archives looking for the data, but we faced a number of problems. First of all, to locate the data. And then when we have the data to actually make the data AI ready for the lack of a better word, very often we will start to look into the data. And that will be horrible. You open the open file and you have some just a bunch of numbers, you have no notation, you don't know what it is. And we get maybe a bit cynical about the value of historical data, because my impression was if you go back to many years, then you start to have something which is so dirty and so unusable that sometimes it's cheaper to reproduce the data. I don't know if you ever experienced something like that. - Certainly, there's a lot of data that perhaps is too noisy and too maybe should be reproduced, but there's so much data that is high quality, that is even, and even that's in the right format, that's the right kind of data, that checks all the boxes that I mentioned before. There's still an incredible amount of that. So even if we set aside, yes, of course, there's gonna be cases where it's difficult to get the data and there's all these reasons, all these barriers, there's still a tremendous amount of potentially low hanging fruit to run some of data that is there and ready in high quality and usable and all of that. So there's no, there isn't really a lack of data. And even of that data that is usable and has all the right, checks all the right boxes, we're using very little of that. And by very little, vanishingly small amounts of it, right? 'Cause again, you have your millions of data points and you have summarized it into seven or two, these sorts of things. So we're making, even with that issue, we're still making use of the tiniest fraction of what we have, right? And then also when, let's say you do have a case where you assembled all this data, you trained your model and it works well, right? And now you have a useful model. Where does that model go? How is that being accessed utilized? That's a big issue too. That's a big problem. Often it's, I'm gonna email the person to build the model and ask them to run it for me and send me the results. Email me the results. This is not gonna scale. And there's certainly, that's the, there are better solutions than that clearly, but there's not a way right now to just say, I want all the world's models to be run on this and wherever they're relevant, give me the answer. Most of these methods are just lost because they're just, there's not a way to uniformly access them, use them. And even when you do use them often, they're only useful in a certain, for certain niche cases, certain cases. And you do want them to come up when they're useful, but you don't, you don't want them to come up with other places. So for example, there's a lot of, far more companies will train models on activity. They'll say, here's my, here's a bunch of chemicals that we have in our archive. And we've tested them against a protein and here are the ones that are active so they actively inhibit the protein. And we take all that data and now we can predict for a new molecule whether it's going to be active or not. And that can be useful. However, the utility is limited to that specific question. So is it going to be active or is it going to inhibit to this idea of activity? If you have someone who says, I want to know, I don't even care if it is going to inhibit, change the activity. I just want to know if it's going to stick to it, if it's going to bind to it. Now this is a whole different question, right? Not an entirely different question, but different enough that the model you built that predicts activity may or may not predict binding, right? Because that's a different thing. So just, I just use that as an illustration that it's, that the models are, they're only predicting what it is you train them on. If you train them on one thing, they won't generalize. We have to at least pay very close attention to how you're asking them to generalize, right? So there's a lot of these nuances that are important and they can get lost and they can, they can limit the utility of machine learning and circumstances. I want to go back to something that you said earlier when technologists, data scientists, and data scientists work with scientists, sometimes they complain and say, why haven't the scientists labeled the data, or prepared the data in a certain way? And my experience is that very often, the scientific process in far-morganization is very linear. You have a scientific question. You prepare your experiments, you run the experiment, generate the data, analyze the data, and then you move on. And then you keep the data, if you're lucky, you keep the data because of regulatory stuff. If you're not lucky, then you just lose the data and move on. And of course, now things are changing because we start to think about data as a reusable asset. Or maybe we think about that. I know a lot of scientists who don't think about data as a reusable asset because they still think about the traditional way of doing research, and not the way that this lab in the loop concept that start in a model, the model informs the experiment, then the experiment generates data that feeds back into the model. And of course, if you want to do that, you need to have the data prepared in a certain way. Going back to something very fundamental. What do you think in pharmaceutical organization is the biggest obstacle to get those different people who have different background, different language who actually work together, the technologies with data scientists and virtual life scientists? I think the technology needs to keep up, needs to solve this problem. I think this is really a problem that I really could be solved by an algorithm, right? Because if you think about web search before Google came around, basically invented PageRank, which is the algorithm that made Google, we're going to use the structuring internet to figure out what's relevant. Before that, there were all kinds of things on the internet that you couldn't find that no one knew about. And so the motivation to put things there or to look there was really reduced compared to now when you can find everything. Now everyone is-- people are going to put things there, and people are going to look there. And you're just going to get more value. So I think also with the scientific enterprise, if we can demonstrate-- here's an algorithm and an approach that when you generate some data, its utility will go on and it will be reusable. You don't have to jump through extra special hoops to annotate it and generate in certain ways, otherwise it will be lost. Something where the algorithm can do most of the heavy lifting so that the scientists does what they do, they generate the data as they do it. The algorithm or the approach takes care of, read the reusability. And that's what I think we've hit on here. That it is so different and so impactful is an algorithm. And again, this is-- we didn't really invent it. So this is Google showed this works for internet search. And many others have shown that this kind of works for internet search. All we're doing is repurposing and saying, it also works for data analysis. You can point this at data. And you can now suddenly reuse incredible amounts of data that have been produced. So that's really the-- so I think the onus is on the algorithms. Because clearly we've spent literally decades yelling at scientists or burying the scientists for, again, the list of not enough data, good enough data and not annotated enough not the right type of data on and on. There's only so much you could do to change the scientific enterprise itself. This is complex stuff that scientists are measuring and capturing. There's a reason why it's been so hard to make it uniform, especially in this space in biology and drug discovery. So that's why we're so excited that we think there isn't any answer, right? From another field, an internet search, there's an algorithmic approach that really addresses this that can unleash just an unfathomable amount of data that we've already collected that's sitting there way that could be reused if only we could figure out how to reuse it. And now we have a method that can reuse it. So it's a good segue to the next part. Because we talked about the relation between technologies, data scientists, with the scientists. But I think we can look at it from the other way. Scientists, our scientists think about technology, think about AI. And then obviously it's everywhere at the moment. When was it, I think earlier this week? There was NVIDIA GTC, the big show, the big annual show from NVIDIA, where a significant part was on life science. Life science is application. There was health care, but also hardcore drug guarantee. Be a more after the rubber, tighten the chips and all those things. And I can see that because it's interesting to see the focus and the investment of really large tech company, trillion dollar company in drug guarantee, in making an impact in drug guarantee. That being said, there is a lot of skepticism. When you talk to scientists, so you will have some of the scientists that are obviously interested, there are fans of AI. And then you have the exact opposite, really skeptical. I don't really see a lot that are in Vermidol and I see the good, the bad and have a little bit more moderate about AI. I think it's one extreme of the other. So talking about your company, because you play in this area. So how do you ensure that the AI driven inside are both reliable and scientifically rigorous? Because that was my experience working within pharmaceutical company when I was introducing tools to some of my colleagues. You just have one chance. If when you do a demo, things is not really accurate. We don't get the thing that they expect. Then very often people will say, you know what, come back in a couple of years when your thing actually works. Yeah. Yeah, there's a lot of skepticism. And I think it's justified to be honest. I'm skeptical knowing people who understand machine learning, right? Know the limitations and know what kinds of data you need for those approaches to work. And the conditions in the context in which they're successful. And so if you then go into the team of biologists or chemists and you tell them, OK, this is going to solve all your problems. It really depends on what the problem is. And it really to fit into machine learning, it has to have certain very specific properties. So if you walk in the door with the example I talked about before, you've got a model and it predicts activity inhibition of your molecule, right? And you go in and say, OK, I'm going to sell you this tool. now it's going to make predictions for your bad activity. You have to be really clear that maybe they want to know binding actually the scientists is I want to know binding activities related to violence not quite the same thing right so there's all these headget and all this Heging and then the question is what was your training set because maybe we have brand new chemistry nothing like it in your training set now the model is not going to work either. So now your model works in very a very narrow domain that doesn't capture what the violence is trying to do right so either the technologist has to go in there aware of that and not over sell in which case. No one gets all that excited about it or they over sell and over step the actual utility and they get called out and they caught because the comments are the biologist like no I wanted binding note my chemistry is two different for you and you should have told me that wasn't going to work and there's all these everyone's I think that's what's leading to the dynamic skepticism and I think it's real. What we've done and this is the reason for our approach so I've been frustrated and people scientists have been frustrated by that dynamic so we're doing something entirely different so you're not making predictions for example strictly speaking our platform is just finding things it's finding data and it's just returning for your question. Here is where there's data the most data here are the kinds of things that data most points to most converges on and here is that data for you to interpret so if you're interested in binding and you look at the data and it's binding data or it's not kind of data right some other kind of data what to do with that. The details are important the transparency is absolutely critical and that's what we're providing so we're just we are allowing scientists basically giving scientists superpowers giving them a capability to find things that they otherwise couldn't find and once they find them because we're providing all the details they can evaluate it they know what to do with it right so it's a very different dynamic that's much healthier that's much aligned more aligned with the scientific enterprise and what scientists need. And so I think once they get a feel for this approach a lot of the skepticism is going to melt away right because you can see oh this isn't someone coming to us and saying here's a prediction and look at the p value the significance the confidence whatever measure it is which know which we don't really know what to do with instead it will be here is something we discover here's something we found. And here's the data the experimental data that we found that supports this conclusion they know what to do with that and I think when they start seeing that and seeing that all that data came from all these different places that maybe they were theoretically they could have been in the seminar read the paper that generated that data point but there's no way they're going to remember it we're assembling it all saying here it is again data that may not have been in the text even not even mentioned in any of the summaries or any of the text it was just buried in the actual raw data we're surfacing all that saying here's where the data points actually and here's the data so it's a very it's a healthier dynamic and that's a benefit for the scientific enterprise. Is there any real world example that that you ought to leave it to share with us where you worked with with customers and help them uncover some critical inside of things that that will clearly not have found without the help of your solutions. Yeah, absolutely so we've been working we've been around for almost eight years now we've worked with all kinds of companies about 40 more than 40 companies about talking about companies. Most of what we do with the companies we can't really share unfortunately but this last year we started working more with academia and those are things more things that we can talk about and we do have a pre print that we just released that has a number of examples in there as well explicit examples but for one example that we've been talking about so we collaborate with a group in the company. We collaborate with a group in University Michigan, Johnny Suckston's lab at University Michigan and they have they run phenotyp extremes which just means that they look for they set up experiments and they cast lots and lots of compounds and chemicals and they ask can I find some that have the effect on my biology that I care about so. I try each one out and see if it does the right thing to myself that indicates potential therapy right so now you've got a bunch of compounds or chemicals that have interesting biology maybe someday they could be a therapeutic because they seem to do the right thing to the biology but the problem is that we don't know how they work we just know put them on the cells and looks great but no idea how they work and it's really important to know how they work. So we were working with them and taking their compounds that are hits in these kinds of screen so one of them focus on liver biology so one of these compounds seems to be promising as a way to treat fatty liver works and mice and but again they don't know how it works. So they came to us we put it in our system was that these compounds that the chemical structure of these compounds is similar to these other compounds here that we found out in the world and in all these databases and those compounds are known to inhibit certain targets. So the DPP4 was one of them and one of these HSD targets is hydrochistero desaturates targets and they tested both of them actually the first one DPP4 confirmed inhibition of that target so one of the mechanisms that we proposed they confirmed and they presented that data in a conference in October in London the L. Ray conference. Recently we got word from them that they confirmed the other one as well the other mechanism that we proposed as well and that's work that we're working on publishing with them so that that will be coming out hopefully in the near future. So there's so we do we've done a lot of work around mechanism ultimately that's I think what where we've had the most impact is we can use the world's data to try to understand mechanism so mechanism of disease how does this disease work. What are our potential therapeutic approaches that we could use to address this disease therapeutic mechanism here's a drug that works or we think it will work and we have some evidence that it treats the disease or helps with the disease how does it work. So our system can help identify mechanisms of therapy therapeutic response also then that leads also to who should be given to because maybe it will work for some of the patients than the others we do a lot of work in precision medicine where identified based on understanding how it works exactly. Understanding which patients will respond to which ones won't and then finally safety what are first of all if you see a safety she can we understand it so we can fix it or avoid it or not have it in the future so safety assessment mechanism of toxicity mechanism of adverse events so we address that as well and ultimately I think all this leads to more effective precision medicine more effective. The business discovery of cures right because when we discover a new therapeutic mechanism that's what we're doing and also more safer medicines right so fewer side effects not giving it to the patient who's going to have a diverse reaction. So when we had chat last time you shared with me a paper so is it this paper but you're talking about the framework for autonomous I driven drug discovery. That's right that's our pre-prepared yes yeah yeah so we will share the link together with with a podcast so you talk about a number of things like the complexity of the data and complexity of linking the different concept so maybe can you walk us through you talk a little bit about it but can you walk us through some of the key challenges when people are faced with data of the load like that. Yeah so I think reading these papers and being attending scientific conferences and just seeing the state of flying across the street and knowing we barely touched any of it. It's very frustrating I'm not sure I'm not the only one I know that I'm not the only one that's frustrated by. So what's really exciting about this approach is that it really it checks a lot of the boxes as far as a solution so if we wanted if we were to solve this problem. We would want an approach that works on different really diverse data this field has incredibly diverse data and that's not doesn't seem to change each any time soon. The diverse data very complex data noisy that noisy data so it has to be very robust to noise. So it really has to capture all that and we wanted to ultimately give us a concise answer because we can't if we just put in massive amounts data and then our result is a long gene list or comment or whatever kind of list. That's not helpful we need something that says no this is the top thing to look at maybe this is the second thing so it's got to be really concise and then ultimately whatever analysis we do. Needs to be transparent especially in size and for regulatory purposes that we need to know here's the reason and the data that led us to this conclusion. I think those are many of the boxes that we would want to check in a solution to the our data overload and the end that's why we're so excited about this approach it checks everyone of those boxes right and the condition because there's always a top head and the transparency we see the data behind it there's robust to noise because it's consensus space so the noise you did it doesn't agree. So you just look at where there's a green so it checks a lot of those boxes for data overload and maybe we'll get to this out let you respond I hope we get to where this all leads but as far as autonomous research but I will get that in a little bit. Yeah we'll get to that so from what I know Plex has developed a focal graph to tackle the issue of data the load how is next gen knowledge graph help filter and refine the drug target hyper phases. That's really the insight that we had was this idea of a focal graph so this is what it is there's knowledge graphs and a lot of people have worked on knowledge graphs and you've spoken a lot about it. So this idea that you can represent knowledge either graph so things connected to other things is the networks. It's a very flexible robust comprehensive sort of data model sort essentially a way to capture. all kinds of things in an organized way. So lots of people worked on that and as I mentioned Google and internet search has worked, has been able to take large graphs, right? The internet is a graph is a network and it figured out how to find what parts are relevant. And so what we did is we figured out a way to connect them. That's what the focal graph is. So what a focal graph is we take this big knowledge graph which could be just absolutely massive and incredibly diverse. And what we do is we plug out from that a sub graph, a piece of it, that's relevant to your question, to your query, right? So now you go from an absolutely massive graph to a far smaller, probably still big but way smaller sub graph is focal graph. And you've plugged out the bits in such a way that you capture the bits that are most relevant to your question, right? So if you care about chemistry then you plug out the bits that have compounds with a similar structure or things that are related to your query. So now you have, that's what the focal graph is. It's meant to, let's pull out the data that is focused on your question and then further add the easiest centrality algorithms that things like page are these internet search algorithms. Apply them to the focal graph such that now the top ranked member or the thing that the focal graph most agrees on is a punitive answer to your question and that's the concept there. Sometimes when you're searching for something there's a number of data sets that will point you in one direction and then there is this incredibly rare data set that points you in a different direction and that actually really interesting and it might come from a new publication so something that is very recent. And if you look at some of the modern tools like well-elems, well-elems are really good at when something is frequent it will come up because it's a statistical approach but if there is something which is very rare, most probably you won't find it because it's not how they are built. With your approach and I don't know enough about how the adrangs work but it was my impression that to find something you need to have many things pointing at it. So imagine if you have something that is very new, that is very rare but that is very important, will people be able to find it? I don't know if my question makes sense but the best way to consolidate it. No, it makes a lot of sense. I know what you're getting at and it's a great question and maybe it's that if there's something unknown, something rare, understudy that we want to understand and if you have a method that looks for lots of different places all agreeing on an answer, how in the world are you going to find something new because you're looking for what everyone knows, right? So it seems like that seems like a paradox. What breaks it, what makes it the way that we find new things is because keep in mind our approach is actually mostly focused on relatively raw data, right? So the places people have described the information, it's usually it's in the text or or pattern something like that where it's been digested and someone's made a discovery and they've described it. So that data is out there and we certainly have that but that's a tiny fraction of the data that we're capturing. The vast majority of the data we're capturing are these relatively raw data points. So I treated patients or cells or something and I measured all the genes, tens of thousands of genes and these went up and these went down and I changed the dose and these went up and these went down. So really relatively raw data of all different kinds. So it's actually usually or very often at least not known what all that data even points do and it's certainly possible we see this often in which is a good thing that the known answer does often come up at the very top or near the top. So we can rediscover from the underlying raw data some things that are known but we can also discover things that are entirely unknown, right? And it's not and we're discovering because it's consensus in this raw data. So it's not that everyone already knew it and so we're showing that we're actually we're saying that the data knew it in a sense. This raw data it was there and in fact it's the strongest signal. Lots of data people collected all over the place all agrees that this is the answer but there may be not a single paper or patent or any written description that actually captures it. So the data knew it but we didn't know it and that's the way that we're discovering this kind of novelty and things that we think of you mentioned you're like rare disease and things that we think we don't know much about. Given the amount of data that we're not using and that we can now use with this sort of approach. A lot of things that we think are rare we think we don't know anything about it are going to turn out and we're starting to see this it's going to turn out that we know a lot about it we know way more than we thought because in fact it's been measured many times still the same people are measuring the same tens of thousands of genes every time and they're focusing on little bits of it we're going to start seeing that the answer is the things that we think we don't know anything about in fact we know a lot about them and here's and here's the answer and here's the data that speaks to it right that no one knew about so it's really going to change how we think about how much we know there's a lot of cures for diseases that are just sitting in that those data sets that we can now expose and say actually here is the answer here's how we know the answer and obviously there's always more to do so people ask me often about validation certainly there's always more experiments to do but the kind of things that we're producing again are not predictions they are here's the data we found experimental data vast majority experimental data that points to this finding here's the finding here's the data that data is often the same kind of validation data that you might then go off to do to reproduce so it's along the continuum of validation so we're taking we're going from sometimes zero where we think we don't know anything and we're running our system and finding actually we know a lot here's what we know here's that data and if you wanted to do a new experiment here would be the next thing to do because again there's always more to do but we can often go from zero to or we can potentially go from zero to a lot in the paper in the pre-prop we show how to go from zero or next to zero to something right and we run it and show you here's how we can take something where we think we don't know and get and take the first step in a transparent way and clearly we can then run that more and build up cases where potentially we have very compelling experimental validation for brand new discovery sometimes we think that we don't know but we know it's just that we haven't looked at it properly so I was recording an episode a couple of months ago with Ben Silagi who used to work at Roche so he was a senior senior leader in one of the data department at Roche and Roche now is quite well known for having invested a lot in the data and digital transformation so he was telling us the story of how it started and how it started to get interest from from the senior executive who learned the paying for this digital transformation and he was saying they partnered with with the scientists with the therapeutic area leader who had scientific questions that they said okay can you look into our old data and find interesting interesting signals that we overlooked in the past and that's what we did they went back did a lot of data clean up good the data and then when they presented they said they presented they explained the methodology it was during a long day of presentation so it was telling me most of the exactly where on their phone were not really interested until they said we cleaned up the data looked into the data and identified a signal by a marker that we didn't know we had and then he told me okay everyone put their phone down and say what did you find and then he told me that from that point on we never had to ask money again because people were convinced that there was value in going back to world data and looking at it from a different perspective certainly yeah I think that's how you get scientists attention as you show them data right you don't I think there's a lot of understandable skepticism about making predictions about models and predictions right so if you if you instead come in and say so here's our conclusion basically lay out something very similar to what they might see an experimentalist might present their latest findings right that they they give the background is okay when we run this experiment here's the result we get and here's our conclusion based on all the data that we've collected right it's just that in our case we're collecting data that's already been generated so it's not to say we don't need to go back to the bench and keep generating more data but the data that we've already generated is just filled with all kinds of discoveries waiting to happen all kinds of results that many of which are could be quite compelling and it's just they haven't been uncovered before right so once we uncover them it's not there's not no convincing specific needs to be done it's just the standard scientific exercise of here's data what do we think this data begins does it support the conclusions and what more what do we do and and where does this bring us so it's something that scientists are much more comfortable and used to to doing and they see they see the value in immediately is there a type of data sets that is better than over at uncovering hidden connection hidden veterans that's a great question i'm sure there is ultimately i think science we want to see convergence ideally of very different approaches right you'd like to see very different approaches yielding the same answer. So you ultimately would really like a diversity of data that all agree ultimately on the answer. So it's, I think some data types are going to be more valuable than others have better information content. A system like this actually, you can, I don't know if it's run it in reverse, but you can certainly do a retrospective analysis. You can say, "Can we run the system?" And we found the following really exciting finding or many findings. And we can say, "What, which datasets contributed the most to those findings?" Right on the one hand. And on the other hand, you can also say, "What datasets, if we were to add them, would be most complimentary for what we already have?" Right. And there are certainly systematic ways to ask both of those questions of the system. And by doing that, we can say, "Let's generate more of the dataset that was most valuable, and let's also generate data that fills in our gaps." And we can use the system to define both of those very, very systematically. Because not only as far as efficacy, we can throw in a cost piece as well. We could say, "Which, what data type has the best value for the buck?" Right. Most cost effective. That not only is information rich, it's also cheap. Right. What's the best way there? And then that enables the system to get even better, right. That rapid pace. Okay. So let's look now and look forward, look into the future. So what are some of the key trends that you foresee in the industry in the next couple of years? And obviously particularly in the intersection of tech, data, and life sciences applications? We've also seen large language models, right, just on the scene. And just it's incredible what they can do. I think it took most people by surprise the capability. And so there's just tremendous excitement there. And what we're excited about is that technology is incredibly synergistic with what we're doing. Right. So if you're looking at large language models, and there's a lot of excitement about agents, for example, or some or retrieval augmented generation, these kind of rad approaches where you have the model, but then it goes out and it asks questions on search as it uses tools or at these agents. So we've seen that these models can be tremendously augmented. For example, if they run an internet search and they find what's on the internet and incorporate that, what what Plex is in this approach, as focal graph approach, is essentially an agent. It's a search agent that finds that can search biomedical data in a unique way, as we see it. And so when you combine that with large language models, now you can have, we've already shown that you could take a large language model and feed it a prompt and say, here's how focal graphs work or Plex works. Here's the strategies you might want to use. And a large language model can actually plan a research strategy using this sort of approach. So we can plan the strategy. We also have shown that given that the prompt with the background, it can interpret the results as well. So you can run your search and you find this stuff and it can look at it and say, here's what's promising in this results for carrying cancer. Here's the key summary. And then in between that, it's pretty trivial to have it also just run the search. So now you have all the pieces. It can plan a research program. It can run searches. It can interpret searches. Could even adjust its plan based on the result and run another search and another search. And this can be scaled. So we've already covered that this kind of system can find novel discoveries, can make novel discoveries transparently. And the planning also of large language model is transparent. So if you put this all together, and this is what I think is really mind blowing about this whole picture, you put that all together, you have a largely autonomous or at least partially, but a largely autonomous scientist that can plan research programs, can execute them, can do this at scale and can deliver novel findings that no one knows. And the data, the experimental data that supports it. So again, not a suggestion based on papers is right, but actually findings based on the data that we've collected and provide the actual data in all the details. And the methods it used, right, what the research program was, the strategy is all these things. So it really, I think it foreshadows a potential truly autonomous AI scientist, which I think is incredibly exciting. People are talking about this, debating whether that's possible and when and how that might look, is the best that I can tell, we've assembled all these pieces, in fact, in our paper, we show it working, at least a little bit, we show it running for 10 minutes or so, right. And finding new things, finding something new, actually a few new things, the cancer space. So I find that tremendously exciting. I think that's really going to have a massive impact on this field. It's going to change the way we approach research. Now scientists are going to be utilizing this kind of tool, this kind of approach, supervising, guiding it, optimizing it, refining it, making sure that it's finding things that are useful, that make sense that pass regulatory muster, right, which is where the transparency is really critical. And I think it's going to really enable a revolution in this space. And so that's the big message we're trying to get out there, to get people talking about it, because I think people don't have them put all the pieces together yet. And I think all the pieces are there to put together, that's trying to tell that story. Are you referring specifically to the paper that was published by a research team at Google, because I think we called it an AI scientist. Yeah, it's called it a Google call it the AI co-scientist. Yeah, AI co-scientist. Yeah, absolutely. And I would agree that's a very exciting thing. And I think the other thing that I find quite exciting, and you talk a bit about it, it's the bridge between the virtual and the physical world. There is all the planning. And then you connect that to automation, to lab automation, and you execute it. And then eventually you get the data back into the system in the semi or fully autonomous way. And that, you know, that sounds like science fiction, but actually it's working now, which is mind-blowing. Yes, oh, it's incredible. That certainly can be integrated so that now you've taken it as far as you can with existing data, which could be really far. I think that's going to really surprise people how far we can go with the data we already have. But there's always more to do, right? So now you get to the very end of what we can do with the data we have. So now the system says, I think what data will give us the furthest that will continue. The success continues to discover it. Tell us what we still don't know. And it will feedback and say, this is the kind of experiment we need to run now. And potentially even that could be automated. It could be the kind of experiment that could be automated. So that gets run and now the data that will push us further ahead gets generated automatically. To close this conversation, obviously, this is, this field is developing super fast. I spend hours every day trying to stay up to date. And obviously, even if I do that, I'm missing tons of information every day. Are there anything that you would recommend that to people who want to stay up to date on what's going on? Things to read, people to follow any other suggestions? That's a great question. I don't know. I'm having trouble staying on top of it. I think everybody is. That's really good to that saying that maybe the filter is use, if you're a scientist, you can't drop your principles, right? So you need to know you have, there's certain things that science requires. And so you keep that in mind and that becomes a filter. So there could be a lot of things out there that are just so lack of transparency and these kinds of things. And there's this current skepticism. I think, I think a lot of scientists are pushing it away until they see what they need. They see the right level of transparency and things like that. And I think it's okay to use that kind of thing as a filter, especially if you're going to be the one consuming it at the end. Obviously, if you're actively working in the field, then you need to be aware of that gap where there's a limitation of your technology and the need. Right? So you can work on it and refine it and try to turn it into something that's useful. But I think as a consequence, if you want to be an important consumer of AI, just be clear about what it is required, what the technology is actually delivering. What that gap is, be clear about it and keep an eye on places where the gap is closing. Right? Approaches that are getting closer to what is required. And maybe on the other side, also some of your requirements, be careful that are they real requirements? Right? Do they really need all of these things? Or perhaps there are some things that this technology is delivering that really are useful. I just need to change my thinking a bit. So it's probably going to require flexibility on both sides of the technologists as well as the consumers of the technology. There's definitely a little bit of meeting in the middle. But I don't know the answer to how do you avoid the overload? For a lot of people, it's you throw it all into an LM and have it summarize it for you. So I'm sure that's part of it. And I certainly use that as much as I possibly can. But yeah, there's a tremendous amount of stuff coming in at us right now, which is very exciting, but also pretty overwhelming. So last but not least, if people want to know more about Plex and the latest news, do we follow you on LinkedIn? Where do we go? Yeah, LinkedIn certainly Plex Research is our company plexershurch.com, is our website. If you want to dig into it, so the preprint that we have, we have this probably on our website. So that AI framework paper will link it as soon as we'll go. So that's a great place to take a look. And I would encourage you, if you do go there, the supplemental files are worth looking at. One of them is a transcript of asking the system with the large language model, hey, can you find a new way to cure cancer for me? And it's a certain type of cancer. We're a little bit specific, but it's a really very pie in the sky request. And the system goes off and does and. tells you what it's doing and tells you precisely what it found and it makes demonstrable progress. Even if it's only the first step, it makes demonstrable progress and verifiable progress towards that. So I think that's very exciting. And then even the other supplement is comparing if you have a large language model without the sort of approach versus with it, the difference is just really pretty amazing. So to have a look at that and just I'm curious about beyond biologists and chemists and people on this space, I hope it's accessible to a wider audience. It was intended to be that way. And if you're interested in having a look and seeing if it resonates with you, if you if some of these ideas really catch your eye, catch your interest, please have a look and let me know. Yeah, thanks. I think that was incredibly exciting, incredible development and says any more to come. So thanks a lot for the discussion. Fantastic. I enjoyed it very much.

Podcast Summary

Key Points:

  1. Doug Selinger founded Plex Research to develop an AI-driven platform inspired by search engine algorithms, enabling scientists to analyze large, disparate datasets (e.g., gene expression, proteomics) for drug discovery.
  2. The platform prioritizes transparency by giving scientists direct access to vendor-line data, making AI insights explainable and actionable.
  3. Selinger emphasizes that algorithms should handle data "as is," rather than requiring scientists to generate perfectly annotated data, addressing common frustrations with data silos and underutilized historical data.
  4. Plex Research aims to uncover hidden biological connections by analyzing massive chemical biology and omics data, supporting target identification, biomarker discovery, safety assessment, and precision oncology.
  5. Selinger notes that while AI has potential, scientists often face skepticism due to past failures, and successful implementation requires reliable, rigorous tools that can prove their value quickly.

Summary:

In this podcast episode, Douglas Selinger, founder of Plex Research, discusses how AI is transforming drug discovery by leveraging vast, often underutilized datasets. He explains that his platform, inspired by Google’s search algorithms, integrates diverse data types—such as gene expression profiles, proteomics, and metabolomics—to reveal hidden biological connections. Selinger highlights a key frustration: scientists are often blamed for generating messy or poorly annotated data, but he argues that algorithms should be robust enough to handle data as it exists.

This approach, he says, unlocks the potential of historical data that might otherwise be discarded. Plex Research’s platform emphasizes transparency, allowing scientists to access raw data behind AI insights, ensuring they are both explainable and actionable. Selinger also addresses skepticism about AI in drug discovery, noting that past failures have made scientists wary.

He stresses that successful AI tools must be reliable and scientifically rigorous, offering immediate value to earn trust. By repurposing algorithmic principles from internet search, Plex Research aims to solve the long-standing problem of data silos and underutilization, ultimately accelerating drug development in areas like target identification, biomarker discovery, and precision oncology. The episode underscores the need for algorithms that adapt to real-world scientific practices, rather than demanding changes from scientists.

FAQs

Plex Research uses an AI platform inspired by search engine algorithms, like Google's PageRank, to integrate vast and disparate datasets—such as small molecule bioactivities and gene expression profiles—to reveal hidden biological connections. This approach prioritizes transparency by giving scientists direct access to vendor line data for explainable and actionable insights.

Doug was frustrated by the persistent focus on scientists to generate cleaner data, when algorithms should handle existing noisy data. He also saw that massive datasets in publications were only minimally extracted for knowledge, wasting vast opportunities.

Plex uses an algorithm robust to noise, similar to how Google handles messy web data, so scientists don't need to specially annotate or clean their data. This allows reuse of existing high-quality data without extra preparation.

The main obstacle is that technology and algorithms need to do the heavy lifting for data reusability, rather than blaming scientists for not labeling data perfectly. Plex's algorithm addresses this by enabling reuse without extra steps.

Plex provides direct access to underlying vendor line data, making AI-driven insights explainable and actionable. This transparency helps build trust with skeptical scientists by showing the evidence behind predictions.

Typical models predict specific outcomes like activity against a protein, but Plex's search-inspired approach generalizes across questions by integrating diverse data types. This avoids narrow model limitations and reveals broader biological connections.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.