Go back

The Patient is Not a Document: Foundation Models for Biomedical AI with Standard BioModel

49m 54s

The Patient is Not a Document: Foundation Models for Biomedical AI with Standard BioModel

In this podcast, Kevin Brown, co-founder of Standard Bio, describes his journey from pure mathematics to building multimodal foundation models for biology. He highlights a pivotal "GPT-3 moment" when he realized that scaling data across modalities—such as neural recordings, imaging, and genomics—could transform drug development and patient care. Unlike large language models that treat patients as text documents, Standard Bio’s approach uses a JEPA-style architecture to map diverse data types (e.g., CT scans, pathology images, EHR notes, whole genome sequences) into a shared latent space. Each modality is processed by its own encoder, allowing the model to handle missing modalities gracefully. The key innovation is temporal prediction: rather than predicting the next word, the model forecasts a patient’s future state in this abstract space, enabling counterfactual reasoning about interventions like treatments or lifestyle changes. This physics-inspired view of patient trajectories helps identify "good" attractors (e.g., longevity) and "bad" attractors (e.g., adverse events), offering a more holistic and dynamic understanding of health. Brown emphasizes the importance of domain expertise and biostatistical rigor in building these models, aiming to create a foundation that facilitates any downstream task, from therapy selection to personalized care.

Transcription

10496 Words, 58226 Characters

English
[MUSIC] Welcome to Data and Biotech, a podcast from Kourtney, where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the World of Biotechnology to understand how they're using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go. Kevin Brown, welcome to the Data and Biotech podcast. No, thanks for us. It's great to be here. Awesome. We'll just to kick us off with you, mine giving us an introduction to your background and what led you to Standard Battle Diode? Yeah. My background is originally in pure math, and then I realized I did not want to be a mathematician. Just for a lot of reasons, although it was quite beautiful. I joined a brain computer interface lab in college. That was absolutely awesome and very cool. I went to grad school for that and got to work on in early versions of what we would call neural link today, invasive neural implants to control robotic arms or virtual versions of those. While I was there, we had to segment some gray and white matter in some brains to prepare for surgeries. And deep learning was just taking off and man, it worked. And so it was just one of these, oh man, the world is different now. Kind of moments when I saw the first few results and I thought this is what I should be spending the rest of my life doing. And some people have described that neuroscience the first time they heard neural recordings. For me, it was the first time seeing civilizations that just worked. So I joined Siemens Health and Ears to work on computer diagnosis, which was absolutely awesome. And that led me all the way to foundation models and ways I think we can unpack. Was there a specific GPT-3 moment where you realized that foundation models were going to be inevitable on biology? Can you walk us through your thought process there? It was really important to record not just from neural recordings, but neural recordings all over the brain. And not just cells, like neurons firing, but also diffusion-weighted imaging, structural MRI, optogenetic stimulation of the brain. All of those things kept adding more and more information, the more modalities and other sort of sensing ways about a system that you could engineer and put together at the same time. While I was doing this brain segmentation, I realized also, oh man, deep learning is going to render some of the interpretable approaches we'd been taking a little bit less performant. And so when I had that oh man moment, and I joined Siemens Health and Ears, I loved grad school to join them because I was like, this is the future. When I was at Siemens, we had scale, and we started to notice that, look, it's less about particularly thinking what are these particular layers that we need? Should we use this activation function or that one? That period was rapidly ending, and we realized, oh man, we're building systems that could be used millions of times a year. We're sitting on a lot of data. What can we do with that? And maybe we should just scale data. Like maybe that's the simplest thing to do. So we did that. We built a federated learning system that, to my knowledge, is still running, which is awesome. That was for computer a diagnosis for lung cancer. But I remember looking at a lung nodule, you know, a CT scan for lung cancer diagnosis. When thinking, man, that can be really cool if we could combine genomic assays or pathology, information, or longitude and electronic health records. All these are relevant and may even bear down on whether or not you expect that to actually be a lung nodule or a false positive or something like that, or indicate specific treatment regimes that would work for a patient or something like that. And the response was, well, look, we make tools for scanners and for radiologists primarily in that group, which is an incredibly noble thing. It drives down the cost of health care and increased quality. It lets it spread globally in ways that are otherwise difficult to do, given the labor shortage of radiologists and the technical expertise required. That's a very awesome thing, but I really wanted to do this multibotal thing. So I joined Bruston-Mire Squibb, where they had one of the world's richest oncology dosets. Rigrously recorded at the cost of billions of dollars through many pivotal, really amazing clinical trials. And that was an absolutely awesome experience, partly because in Pharma, the domain expertise is just higher than almost anywhere else in the world. If you want to talk to a world expert in self-therapy or gene therapy, you put time on the calendar, you call them up, and then you just ask questions, and they're happy to talk. You want to talk to someone who really understands how imaging works in clinical trials or should work or could work in the future. You talk to those folks, likewise for biostatisticians. I was really able to just like, you know, some of the most productive knowledge and merst times of my entire life, including undergrad and grad school, which is being able to walk the halls of BMS. There was a moment, a GBT-Thief moment, which really was literally GBT-Thief moment, where I read the GBT-Thief paper. You know, we've been building data science models in drug development, right? So human data predicting outcomes of clinical trials, let's say, or response to therapy for a patient or something like that. Those were great, but I went, oh man, there's this one figure in the GBT-Thief paper in particular, where it shows the performance on several downstream tasks, like 100, and some of them, you go from like 0% accuracy all the way to like 100 at different scales, like sometimes it takes 6 billion parameters, sometimes 13, sometimes more. And I thought, man, that's going to happen in biology. And we're going to make these multimodal. It's all going to go into one giant foundation model. And we're going to figure out a way to make it work. I don't know exactly how it's going to look, but that's what's going to happen. That was sort of the GBT-Thief moment, and then the rest of, you know, the next couple of years we're figuring out where and how to do that. Yeah, awesome. So with that background, can you give us an introduction to standard model bio and maybe how you think differently about foundation models for biology? Yeah, for sure. So I realized we needed, you know, multiple modalities, and that goes back to my time in grad school. But then I also realized we needed data scale. And there was a third element that really came down to me, you know, that I started to understand when I was at BMS, which was that that domain expertise, that last like 10% was critically important, right? And it's not something that every foundation model can do all of. Like you can't own predicting outcomes in metastatic melanoma and non-small cell lung cancer and early breast cancer and cardiovascular disease and cardiovascular disease if a patient has a previous history of a GLP1 or a diabetic history or something like that. All of those details are critically important. And no one can build a model that does everything, but maybe you can build one that facilitates any downstream task. So I thought where would be the best place to do this? And where would I get the data to do it? So I thought, well, look, maybe I do this at BMS. BMS is a fantastic place with incredible data, but that data like every big farm is data is highly biased. And it should be, it's mostly your drugs, mostly early trials, mostly failed trials and older standard of care, not a lot of your competitors' drugs. And that limits the generalizability of any foundation model that you want to be very general. So I thought, okay, it needs to be more than just these vertically specific things. Yeah, awesome. So, you know, in terms of, you know, how you think about foundation model, the facilitates downstream tasks and enables a variety of different use cases. One of the things I've heard you argue is that the patient is not a document. So can you, you know, unpack that and sort of like give us an understanding of like how you think about the patient mathematically and, you know, why is, you know, thinking about a patient is sort of like a bundle of text, you know, fundamentally limits, limiting. When we first started this, we thought, look, we're gonna use this with, with text, right? We thought we'd use a large language model backbone. You'd be sort of silly to not use an LLM at all, right? For biomedical, yeah, because they've already been trained on like all of PubMed, right? And most of the scientifically published literature for better or worse. And that makes them tremendously powerful tools. And so we thought, well, look, language is a very general space, right? It's almost by construction designed to describe anything. So maybe we should just map everything into it, whether it's radiology images, through radiology reports, whether it's pathology images, through pathology, the initial pathology reports, whether it's, you know, any other molecular report. That was fairly fruitful. It's not that difficult to do to give you an example. We used an existing pathology foundation model published by Bioptimus. We took that, we combined it with a Mama AB model. And in two hours, you can get state of the art for medical image understanding, you know, at the time, on like a single H100, including debugging, right? So we were thinking, oh man, this is the future. But then we realized that when we wanted to incorporate other modalities that weren't as readily mapped to text, we were running into some walls. So my favorite example of this is whole genome sequencing, right? You've got, you know, 3.6 billion base pairs. And you're trying to compress that into text. But most of the text you have that's linked for any of one of those things is like K-ROS positive, you know, EGFR positive, like great, right? You know, ROS-1, okay. But that doesn't really give you a rich enough way. So if you take some of these contrasts of approach or the traditional vision language, or other modality language models, it wasn't really going to work to then feed it into like a language aligned embedding to add to the context to the LLM. So we switched a little bit to a Jepa-style model. And that was one of the first things that like really felt right, like we think, oh man, like we are really respecting the complexity of what it means to be a human patient, right? And then every physician, every clinician knows that a patient is not just what's written down in the notes or even just communicated it around. There is other subtle things. There is, oh, well, we, you know, we assume X, Y, and Z because of their previous history or things that don't get written down or maybe they just get started out loud separately. There are visual things, right? Then an epileptologist can see when they look at EG traces, right? And they don't always, again, make it into the report. And so we thought, wow, this works really well for radiology and pathology for other things. It doesn't. Let's map it into a shared latent space. So we still use language models. Those language models can interpret physician notes or electronic health records, map them into language space. And then that language gets projected into the shared embedding space, likewise for images, likewise for digital pathology. And this actually turns out to be not that difficult because every single AI model, more or less, maps input data to some compressed representation that's a vector. And you just need to find a way to take that vector and map it onto the shared internal vector, whatever. That really made us happen. We thought, well, what's the best way to pre-train this, right? How do we make sure that? that we're getting a fairly good model. Should we use data corruption techniques, like mass language modeling, for instance, or mass image modeling. And we really were inspired by some of Yon Le Coons' work with the job models joined vetting predictive architecture. One of the things we loved about that is it naturally fit into this patient as not a document perspective, where we take a lot of modalities, the imaging, pathology, EKGs. We map them into a shared space. And then you predict what that space should look like at time T plus one, and then T plus two, and then T plus three. And that shows it's almost by definition a trajectory of a patient in this abstract space, of patient space is defined by the different measurements that you could make about that patient. And it's never limited to the ones you maybe have for that patient if you have an EKG grade. It can give them a little more light on where they are in that space. If you don't, that's fine. You kind of marginalize over it. And we felt like this is the first thing that really respected sort of patients. And also, let us think about patients moving temporally. They're not just a snapshot of the document you have at a particular time. It's not just about predicting the next word you've seen in a note or answering a question that may go into a radiology report. It was really about, hey, where here's a patient today, where do we expect them to be? What I'd love to do is just kind of say back to you my understanding of the way the model works. And like it may be like work with you to define some terms just to make sure that you know that everyone's coming along with us. So first of all, like ultimately what we're trying to do is create an embedding for each patient based on their medical history up to that point that captures all of the information from all of the different modalities, all of the different ways that we measure that patient's health throughout their health journey from everything from CT scans to blood work to genomic profile, clinical notes, all that ends up in this single vector, the single list of numbers that says, this is who this patient is at this given point in time. And the Jeppa approach that you're talking about, my understanding is that rather than trying to take an approach where it's like trying to make the patient noisier and have the model predict who that patient is, what you're doing is trying to predict what that embedding will be at some point in the future based on what's happened previously. So it's like you're like stepping through time and saying in 2020, this is how we thought the patient would look in 2021. In 2021, this is how we thought the patient would look in 2022, et cetera, et cetera, up to the current moment where you can then take that embedding and you can do things like predict what that patient might look like in 2027 based on their medical trajectory to date. Am I thinking about that right or would like where? Yeah, exactly. You're thinking about it right. I think one of the things that it affords, once you get into this regime one, you can build anything off of these embeddings that you want. Just like every AI model takes data, it's seen and produces a compressed representation and does something with it, whether it's a classifier head, basically putting logistic regression or stop next head on top of a classifier, whether it's regression to a single number several, or whether it's a diffusion head for producing an image, for instance, in a stable diffusion sum model. We know that that's really almost like Plato that anyone can go build things off of, which is really cool. But the temporal nature of it, what gets as excited is, you know how like in physics sometimes working in the space of differential equations makes things easier, right? When you think about the equations of motion, it's easier to specify them in terms of change over time and space rather than the mapping out the explicit trajectory, right? Integration is kind of hard, but differentiation is kind easy and it's easy to say, hey, look, if I've turned this a little bit, this way, this is where I expect the ball to roll or the planets to move. And I think we're moving. Like I really like this sort of physics very boost analogy, right? But when you think about the patient, like I feel a little bit sometimes like, you know, Kepler looking at the stars and trying to figure out the equations of motion, right? And writing down like in a notebook, oh, here are the stars, we're in this position, that position, this position, and then trying to figure it out. And then once it like sort of clicks, that oh, wow, you can describe these things in these very elegant mathematical ways in terms of like how you expect them to vary over time, everything kind of makes sense, right? And going through classical Newtonian mechanics. And we thought, man, like I want to get to that place for a patient where we say, hey, look, this is where you are, this is where you're going. And then you can think like you can start doing these counterfactual games, right? Would the appropriate statistical rigor, like I'd death that was so drilled into me a BMS, but like I came out with massive respect for the biostatisticians there. And you think about like the difficulties of really designing systems that ascertain what's true. But I love this idea of counterfactual reasoning of like, hey, look, this is the trajectory you're on. And if we do this intervention, whether it's something like exercise, whether it's a GLP1, whether it's a cancer treatment, we expect you to maybe go over here, right? And then you can say, hey, look, not only did I do counterfactual reasoning or I jittered some things to see how likely it was that I was going to go each time with sort of some carless sampling or something like that. You can say, hey, look, am I on the right track? I did this intervention. That means I expect that I'm going to be over here. Am I actually over there? Is this therapy working? And I love this idea of thinking about the good spaces of patient space, right? Where you have longevity expected and things like that. And these sort of like bad attractors, right? Like serious adverse events, right? Or cardiovascular events or like a heart attack or something like that, or like metastatic disease and trying to push a system away from that, right? And you can sort of start thinking, hey, maybe these therapies are basically bumps and maybe they work for a period of time, like a targeted therapy that works great in lung cancer for a few years before you develop resistance or something like that. And I think this is just a way of thinking about it that is just kind of elegant. And also, you know, respects the fact the patients, like they change over time. They have many measurements that matter. They're not just what's written down in one specific document. They're not an ICD-10 code. I want to spend some time on the modalities aspect of this and then kind of return to the temporal nature of it and like how you think about that, those counterfactuals and like how a patient moves through time. You know, my understanding is that you're taking these, you know, very different data types from radiology images to clinical notes and you're sort of treating them differently within the model architecture. And you're utilizing different assumptions in the way that you process each modality in order to get the most information from that, from that particular modality as it's being brought in. And that also there's a way of handling this that means that you don't have to have every modality for every patient in order to effectively model that patient's trajectories. So can you maybe help me understand like how those different modalities are treated? - Yeah, so the different, each modality gets its own specific encoder to begin with, right? So EHRs can be ingested by language models fairly readily. But then likewise, let's say a radiology image, right? We, you know, we use vision transformers in 3D. It takes, you know, really about three lines of code to build them today, right? It's a fairly straightforward thing. Of course, then when you want to do a little more fancy things, you have to think about various data augmentation techniques and things like that. You know, we've, for instance, trained on a million CTs, maybe 1.5 million now, but a million oncology related ones. And that does allow us to sort of encode various kinds of CTs and then put them, you know, into that same, you know, fight space will be coming public with that soon. But like the idea being like, we treat vision, in a way that we know there's good standards for how to treat it. We don't have to reinvent the wheel, wheel there. And if there's a better radiology model, we can also use that too. Likewise, digital pathology, we use other digital pathology models right now. I think the field is like a wash in them. I think it's a really intense and crazy space where there's like all this churn and, you know, one week, you see this $50 million, whatever model. And then in the next week, you see someone train a very similar one for a few thousand. And we thought, you know, what we're going to like just sort of use what other people have produced at this point. And that's totally fine, right? So you can ingest using best practices, digital pathology data, with all the careful thought that it has to go into that, and then map that into that shared, you know, shared latent space. Likewise with an EKG, likewise with, they could be EEGs, it could be brain MRIs, it could be, you know, at chest x-rays, things like that. The important thing is we don't have to develop every single encoder ourselves because I think that's too big of a bite to chew. But we do that in imaging. We do do that in genomics. We are doing that soon in multiomics and other things. And there's a number of reasons for that. The other thing that I love about it is that if you're missing a modality, that's okay, right? And this naturally pairs with the way the clinicians have to operate. Sometimes you don't have a test that you would like to do for a very, you know, a variety of reasons or it's missing or it didn't come at the right time. Or, you know, maybe the endoscopy really didn't get what you want to see. And you have to think about how that information would get built on otherwise, you know, things like that. That means that, you know, you're not totally, you're not totally bound to these very low-scale regimes where you only have really highly paired data for every single sample. I think that's unrealistic and you're not going to get the scale that you need to build the size models that you want to really reason about these things. And I want to maybe spend a little bit of time. Why do we care so much about so many different modalities? All right. What if you just did a vertical model for, okay, here's my radiology model? Here's my pathology model. Here's my EKG model, you know, here's my genomics model. And then I'll just sort of integrate them later. I think that doesn't work for a number of reasons. One, because even at a low level, you get some interesting cross-attentially, you know, feedback across these different things. But two, clinicians know that like, they don't even operate in these intense verticals. There are some silos, but even like a radiologist doesn't just report or produce a report, and that's the end of it. Oftentimes, like at a serious hospital, there will be feedback. You can call them up. You can ask them for clarification. You can double down. You can say, hey, well, the patient was reporting and so I think you take a look. There is that kind of crosstalk. And so it makes sense to build models that also have that communication built in. And so I think that is a future, or one of my favorite examples about this is cardiotoxicity, right, to show just how broad these things can be, cardiotoxin in oncology. So in radiation therapy, you can have adverse events from radiation therapy, especially in the cardiac sense. And you can triage patients, according to how likely you think it's going to be that they're going to have that kind of event. So you can look at pericardial effusion in a CT scan. You can combine that with routine electric cardiograms. You can also combine that with the longitudinal EHR. That would obviously show different kinds of risk factors, whether it's as simple as body mass index or similar, previous obesity, history, smoking status, things like that. Well, all of those are going to be relevant. And clinicians are going to compare all of those things, and why not do it in sort of raw signal space if you're building a model. It also shows how interdisciplinary we know that this model has to be. So the verticals can't even be like oncology, right? Because even in oncology, now you're in cardio, right? And those cardio models are going to get better, even if they're just trained in general, cardio patients. And likewise, you're also going to be in immunology, not only because of immunotherapies and the side effects that you have there, but like, look, there's gastroenterologists that only see oncology patients. There are cardiologists that only see oncology patients, and it's only going to get more complex. And so we didn't really see there be, like, a limiting factor there. And then once you're that wide in the horizontal, like, man, there's no way you're going to own all those application layers. You cannot credibly claim that you are an expert in cardio toxicity for early lung cancer and the progressive metastatic disease in melanoma. Oh, and by the way, just general cardiology, like that's insane. But you can do that first 90% is very similar. Yeah. It makes a lot of sense. And also, you know, like the goal is a foundation model that is the best, most complete representation of the patient that's possible. And there's information in all of the different modalities that can inform that representation of the patient through time. And so to ignore that is to fundamentally restrict yourself from getting that full representation of the patient. Another question I had about the modalities is like, how do you prevent a high signal modality, like imaging from dominating, like, noisier modalities, like genomics? That's a great point. And this actually plays into the temporal aspect. And this is a problem we're actively working on. But if you're an oncology patient, you're going to get several CT scans, most likely. And they're going to be at regular enough intervals. And so CT is a great place to go. That's why that's the other modality that we went into besides longitudinal EHRs because the most regularly paired. And then a baseline, you know, you may have sequencing done, right? But it's not going to be like all the time, right? Likewise, a resection with pathology, very important to baseline. It's not going to happen that the same frequency is CT scan. So there are ways in which you can get better clarity of where our patient's going on the trajectory. And then that can be updated to seeing what you expect to see and the MRI. The other thing that we do, which I think is important, is you have to kind of ground the model a little bit. So you need to either be able to like, you know, reconstruct the signal you would expect to see, right? Like, hey, am I seeing a plausible MRI at this point in time for where I am in embedding space? And I seeing what would be a plausible doctor and noter ICT-DN-CUD at that point in time, or am I seeing an EKG that looks like I expect? That I think is a pretty important thing to make sure that the embedding space stays relevant and it doesn't collapse into sort of repeating sort of trivial embeddings. And that was a major mock for us. - Right, so if I'm understanding what you're describing is like that you're using like a joint objective when you're training the model. You're using like the supervised version of the model where you're trying to predict the next quote unquote token, the next thing that would be in that patient's medical record and then you're also using Jepa's you described as you described earlier. - Yeah, yeah, that looks right. And we were pretty happy, you know, in comparing this to just, let's say, predicting next token or medical history event. We worked with the more excellent catering with an absolutely amazing, you know, data set that combined things from digital pathology, imaging, longitudinal EHR and, you know, I would say EHR plus 'cause it, you know, contained a lot more than you would expect and very temporarily done and just very rigorously curated. And that was an amazing collaboration 'cause that allowed us to test this out, right? And so for model size, even for bigger models, you know, compared to bigger models, it does provide a lift. And interestingly, the bigger lift is not in some of the simple things, but it's in producing like certain kinds of toxicity or progression to specific organ systems. And also it works better in more complex diseases like sarcomas. And I thought that was like pretty cool because I just, I think there's certain things where you're maxed out, you're not gonna get too much more prediction in terms of like predicting outcomes if you already have maxed out sort of what that data's capable of doing. So we were really excited about that. Like that one, when those results came back, I still remember Ershad who really really led that work. He's fantastic. I think he did a lot of it probably while he's still defending his PhD, but I was in India at the time visiting a for a friend's wedding and he called me out. And he got out of bed and he knew he was like dancing on the other side of the world and we were both sort of like jumping up and down just like yeah, this is cool. And it really felt right. There are a variety of reasons that that makes sense to me as well because you know, I still sometimes can't get my right. Mind around is the idea that the model is constructing the target embedding as it's trading the predicted embedding. And so like having that supervised objective like to me, at least makes sense in terms of like helping the target embedding along toward something that the predicted embedding is getting there when it's you know, using that to calculate the loss inside, inside of Jeva. And what I love also a bed that I think makes it possible to do is like look, if I take this additional measurement, am I going to meaningfully change where that embedding would be or what it something else would look like in the future? Is it going to inform what I would see with other modalities? And then that lets you optimize the kinds of tests that you would run, right? Whether that's in clinical trials development and picking the biomarkers that you would want, you can maybe figure out which ones are redundant 'cause they're not changing the expected trajectory of a patient, so I don't do a measurement or whether it's in a system that's eventually clinically facing and that makes me pretty excited. 'Cause you know, everyone likes to say they're right-dark for the right patient at the right time, which I think is cool, but it's also like the right measurement for the right patient at the right time. I think informs all of those decisions. So this was another question that I had about the modalities is like is one of the downstream tasks like figuring out what the information gain would be hypothetically from taking a particular test from a particular modality for a particular patient in order to like improve the embeddings? Yes. Yeah, yeah, yeah, we would love to do that. And like there's all these nuances for doing it, right? So we can like play around with them, we can do these basic assessments and we do, we are actively looking for like really awesome biased ads, collaborators that would love to just take it and do it. You know, I'd also mention like in terms of these downstream tasks, like right now all the models are open source with very permissive commercial licenses, right? Like you can go and download them and you can use them and you can do different various things with them. So like we really want everyone to go take it and do beautiful validations with it and beautiful downstream tasks, right? Like and we'll even help people do it. We worked with one group and you know, the paper's not public yet, but it took them like a couple hours to get a positive result. And then you know, they submitted the paper like a month or two later, which is fantastic. Likewise, you know, we repeated that with some other groups. The other day one of the other like really highlights for me was I just told Cloud Code, hey, like go down with the standard model and apply it to this data set from this paper. I didn't provide any links, right? And it like searched the docs, pulled it down, it didn't require any intervention. And I was like, okay, that's cool, right? I think this is now I feel like we're really spying up biometically in like a meaningful way. I made me happy, yeah. - The tools that are available to fine tune, to fine tune foundation models, you know, across tasks have just improved dramatically and that you know, what that adds up doing is just expanding the potential list of end users who could think of a problem that, you know, at a time aware patient embedding could, you know, help to inform. I'd love to get to some of those, to some of the tasks that you think would be most appropriate later, but you know, I think one of the things that I'd love to understand a little bit beforehand, when you think about the temporal nature of the model, how do you separate something like disease progression? Like the model is picking up on this patient is likely to, you know, have some progression and some, in some disease that they're experiencing versus there's some treatment effect that, that, you know, happened earlier that is then, you know, altering, altering the patient's trajectory. I was just having trouble thinking of like, if the tumor is shrinking after chemo, is the model learning like, chemo works, or is the model learning like this type of tumor was, was, was, or, regressing or like, how do you think about like, counterfactuals and that kind of, we think deeply enough about that to know that we should have other people think even more deeply. And, and, and part of that is, you know, look like causal data analysis is really hard and understanding the heterogeneous, you know, treatment effect, like a treatment effect for a specific patient is really difficult because you can't go back and give the same patient a different treatment and see which one worked better, you know, you can't go back in time. So that makes it really different and there are very good statistical techniques to do that. They do require a lot of nuance. I think it does permit that kind of analysis, but you have to be like, going, again, going back to the BMS days, like, you got to be really careful, or you will be yourself down this road that is just like not statistically, sound. Another thing about the time, about the time aspect of it. Like we talk about time in terms of like T minus 1, T and T plus 1, but sometimes you're trying to predict what the, what the patient's going to look like in two years and sometimes you're trying to predict what the patient's going to look like in six months for people who are sort of like downstream users of a model like this that are trying to predict the next thing that's happening. How do you, how do you advise people to think about like time and the context of the model? Yeah, that's a great point. So we would love to make it very time agnostic from very small scales to very high scales. Where I would love to see it goes, hey, look, in the next hour, what should I do with this patient all the way to like in the next year, what should I do with this patient? Or people build systems that do that? I do think that will be difficult. Like we're not in the like real-time sense yet of like, hey, you should do this test. But like there are scenarios where you would want to be able to make those decisions, right? With a patient that's like likely to go into status and that's going to change like different treatment things that you could do or has a sepsis likelihood or something like that. But it's like, I think one place that could be cool like applied in a cool way. And there was some actually like LLN based work at NYU Langan that did a lot of this, which I thought was really awesome. Both senior PIs on that work, King and Cho and Eric Orman, they're absolutely fantastic researchers. They did implement in real time being able to predict things like mortality and in, perspective, shadow mode, which I thought was great. I would love to see it used to like reduce alarm fatigue or something like that, right? Like, hey, look, like what is the most important thing to alert a clinical team to for this patient at this point in time? Or what does the test need to do to avoid those things? And how can I queue up the the racks for the things that will need to be done? We're not there yet. But I would love to see that like, because you're right, like patient relevant time skills can go from seconds, right? And the stroke to like years in terms of like exercise intervention. So even though you can't know what the like what the time what like what the time scale is that you're operating on, it does allow you to do things like predict T plus one, see what the embedding looks like and predict T plus two, predict T plus three, predict T plus four and see like in the absence of intervention, this is what the patient's trajectory looks like through amorphous time. And then also sort of inject a study or an intervention in the like in the middle of it and see how that influences sort of the trajectory of that patient. Yeah, they could have out there. Yeah, that's right. Yeah. Let's go back to foundation models versus vertical models for a minute. So like I know that you've like you talked a little bit earlier about how medicine is not naturally narrow to these disciplines that like there's information from each discipline and each modality within each each discipline that sort of feeds into the overall picture of the patient. Can you just talk about, you know, do you have any like what evidence out there supports the claim that like narrow models are more brittle or that the approach that you're taking for these brought through these broader foundation models is, you know, definitely the right way to do it. So formal evidence like I wouldn't go so far as to say that like it's not like that's not a registered, you know, a registered study for it. But a lot of it is just sort of reasoning through it from first principles. Like I think that we know that multiple modalities are all of our patients or we wouldn't have tests from those modalities. I think the question is what is the best point to integrate them and the way they're integrated right now is largely text and then sometimes follow up text, right, a radiology report that gets read a few words that describe a genomic assay, right, EGFR positive or whatever. Those are very valuable today, but integrating them at the end of the text is like, you know, by the time you get something all the way down to text, you've reduced the content. Like even in a normal LLM, right, like predicting the next token, one token has less information than the embedding that you use to predict that token. And so it makes sense to communicate across these modalities and embeddings in one way, shape, or form because there's just so much more richness there. You can only ever destroy information by converting it into into text, even if it makes it more legible for a human. And then the question is where should that communication happen? Should you do it all the way at the end? Like with the vector that you would use to predict the next token? Should you do it earlier? If you do it earlier, do you get benefits? And that's what hasn't really been rigorously assessed yet. I mean, I know there's some papers and research that looks at these levels of fusion, but I think some of it hasn't really explored the scale at which you really could or should be doing these things. And so I think that, you know, it's going to make sense to have multimodal models. Now, medicine does get siloed, right? There are tracks, obviously, and there need to be for reasons of education. But I do think in the future, and I don't know when, but I do think that some of the communication that happens right now by text, either by physicians or by agents, right? Or like you could have a radiology to text report generator. And those reports then get fed either to another human or an agent that is interpreting all of that text and combining it with text from the genomic essay. I think in the future, it's not going to be an agent or a human at the top just pulling the text. I think the communication is going to happen implicitly at a lower level from modality and modality, from a model that is sort of naturally communicating because there are embeddings that are talking to each other from radiology to pathology, to genomics, to general patient information. And the human will definitely always be in the loop, but I just think that the communication can happen at a lower level because text is just such a loss of medium, right? Right. So the argument here is basically that every small piece of information that's gathered throughout the patient journey should be treated as like gold, as like really, as like really valuable and that this embedding space that you're constructing across modalities is designed to capture as much of that information as possible and not lose any of it through the ways that humans communicate with each other inside of hospitals, you know, inside of EHR, EHR systems, and that the EHR systems could potentially benefit from having systems that reference that's, you know, condensed representation of all that information. Am I? Yeah, no, that's right. Because like, and this isn't, you know, really like new when you talk to like a radiologist, like they know there are things about that scan that they're looking at that don't make it into the radiology report for a number of reasons, right? And like it's there, they can see it, right? They can look at a tumor margin and maybe think, "Hmm, that doesn't look necessarily as good as I would want, but like how do you quantify looks icky?" Right? Like, you know what I mean? So, but it's there. And likewise, with, you know, digital pathology, sometimes it's not just about counting up, you know, like a PE01 measurement. It's a little bit more nuanced and some of that nuance doesn't get put in for a number of reasons, but it is there and it can be picked up. And I think that the labels for it are difficult, but if you have a self-supervised task, which is like predict where this patient's going, that's highly patient relevant, so it's meaningful. And it's also I think naturally pulls out those things that matter for, for that patient, even if they don't get written down explicitly. Yeah, I think that makes a lot of sense. And also it highlights for me the way that our, our healthcare system is sort of like trying to get everything into discrete space, you know, everything gets a code, everything gets a category. It's a binary like you're either diagnosed with the thing or you're not diagnosed with the thing, but that the embedding representation can capture like a more continuous journey of, you know, what those different radiology images look like that say, you know, we're moving in the direction of a diagnosis that looks like this, even if the it is not diagnosable at time t minus three. Yeah, yeah, that's right. I mean, I completely agree. I want to talk a little bit about the training data that goes into this. Also, how that plays into like if a researcher is listening who is sitting on a bunch of patient data, you know, how does the way that the data is prepared to go into this model relate to the way that relate to the way that, you know, somebody who has a lot of patient data and wants to get these patient embedding representations out, how would they, you know, interact with the model in order to do that? Yeah, that's a great question. So it takes way less data normalization or curation than it used to, right? In fact, when we're based to the table, we back it out into text anyway, because it's easier for the LLN to ingest. I think, so don't be, you know, scared away by thinking, oh, no, like I'm going to need to like spend hours and hours and hours like making sure all of my columns line up to whatever I think that the columns were the standard model it's using. I don't think that's maybe the best way to think about it because it's going to go back into text anyway. I think anytime you have those raw modalities, it's fine, there's relatively limited amount of pre-processing you need to do. And I think that makes it really good as long as the data is accurate and high fidelity. And that part, of course, still needs to be done because sometimes things get mislabeled, or something, so you're going to ask that we've seen people accidentally write micrograms instead of milligrams. And those are obvious to correct when you see them, but like the model may not know. That, so there's, you know, I need to be reviewed, but broadly speaking, it's not that difficult. And we're happy to help anyone do it, but, you know, it's not, not crazy. And is there any concern either on your part or do you think there should be concern on the part of potential users of these models? The data that you've trained the model on is from a fundamentally different distribution than the data they might be feeding into it, like some data shift happening. That is a fantastic concern, right? So right now we are on ecology forward, although we're moving rapidly into immunology and cardiology as well. That being said, yeah, if you went and downloaded this and said, hey, I'm going to use this model that is really good at oncology to predict, you know, neurodegenerative outcomes in my Alzheimer's disease trial. Well, you're going to have some issues without like fine-tuning that model on on some Alzheimer's, you know, Alzheimer's data, right? I think that that would be one thing. I think it goes to a broader point, right? Any foundation model is a compressed representation of the data was trained on, right? Like that's kind of by definition. And so we want to be sort of this latent super, you know, like pseudo data. layer so that you don't have to go out and buy all of that data necessarily, right? In order to train your model, you don't need to go buy thousands and thousands of CT scans or millions or whatever. And I think that makes it exciting or easy for people to get up and use. But if you're trying to use that model, no matter how good the model is, if it wasn't trained on cardio data, it's not going to get a cardio. If it's a task that is not in distribution for the data that you originally brought in when you trained it, you should be thinking about how can I fine tune this model and that evaluate it, which brings me to the question of, you know, what are some of the ways that you make that you evaluate the quality of the patient embeddings that you're producing? And how do you interpret how do you think about interpreting how good the model is at different downstream tasks that it might be asked to do? So that's one of the core questions, which I think actually doesn't get answered enough, right? So we're at this period where everyone is like, oh, I want to build a foundation model. And so they train a model and sometimes it's only evaluated on one thing or two things. And that's not really what you would necessarily want or need. You really need to evaluate on lots of things, but the problem and the reason why people don't do that is not that they're lazy. It's that they don't have the data, right? And so you need to go to other institutions and say, hey, look, like, can you evaluate this on your use case on your patient population? I think taking one step back, if I'm a patient and I want a model that's powering some AI that's making decisions for me, what makes it good? Right? Well, okay, we say the embeddings that the application layer is using are good. Fine. But how do we know that? It doesn't really matter to a patient how well it works on average. It matters how well it works for patients like them, right? And this is a fundamental problem. Everyone has this problem. You know, when you go to a clinical trial, you have a particular population, but that may not be what it's like to be applied for a lifelong spoke smoker who grew up in, I don't know, Greenville, Alabama or New York City, right? And we know that the socioeconomics and various demographic things have real effects in patient populations and affect the way that things are applied, whether it's at the drug level or the AI level. We need to evaluate these things very locally. You know, I'm not the only person who said this. I think Nameshot that Stanford has been saying things like this, like local evaluations are good. And they shouldn't just have huge benchmarks that are like, hey, this is the lung cancer benchmark. This is the, you know, a senior special lung disease benchmark. We need ones that are like really specific to the patient populations that they serve, which is why we're happy to give the models to different people to go evaluate it. I can actually think that needs to be done more because you see these patients and it's like, well, we trained it on this data. It was really good. We got good results. Huge scientific insight. But we really only know if it works to this one institution. You know, like it needs to be validated on more places, I think. Yeah. And, but practically speaking, you know, based just based on my understanding, it looks like use the data that you have to get to get embeddings out of the model, either through inference or through fine tuning and then, and then inference. And then you, you know, you use some sort of prediction head on top of those embeddings. It's like a, it's, you know, it could be as simple as like a regression on the, on the different features in the embeddings to predict a particular outcome. And you say, you know, how predictive is this, like linear model or more complicated model of the outcome that you care about. And that's sort of, you know, how one would think about evaluating the quality of the underlying embedding representations. Am I thinking about that right? He's doing something that's useful. Like in oncology, let's say I'm doing time to event. If I put a Cox proportional hazards head on top of this model that takes these embeddings and spits out, you know, the hazards. Can I use that to get a higher concordance index or time dependent AUC or whatever the appropriate metric for that particular question is, then some other model, right? And that would indicate whether it's useful. Right. Yeah, that makes a lot of sense. That makes a lot of sense. So it from a use case perspective, can you just highlight some of the ways that your open-source models have been used in the wild so far? Yeah, yeah. So ones like, you know, like early onset pancreatic cancer ones, this, you know, thinking about cardiovascular toxicity. The other is toxicity for general treatment for, you know, whether it's sarcomas to pancreatic cancer. It's a non-small cell lung cancer. Those are the first places where we've seen a lot of the application. Some things that are relevant to clinical development, right? Can you predict when a patient is going to have a line of therapy transfer that would make them eligible for a clinical trial? Can you predict whether or not they're likely to meet the inclusion exclusion criteria for that trial at a specific time in the future rather than just right now? So that eventually, HCPs can take the model or applications built on that model and then start those conversations early because you expect the patient to go in a certain trajectory, you can start figuring out what the best trial for them might be, for instance. Or eventually, you know, in the, like one of my favorite places is thinking about metastatic disease for oncology and can you predict certain kinds of adverse events? I think it's going to be a huge thing, right? And I think it's going to be fun. Yeah, the other thing that's occurred to me as we've, as we've been talking about the embedding space is just that the embeddings are so flexible in terms of the things that you can do with it. Yes, you can feed it to a deep learning model to predict something. Yes, you can, you know, run a regression on them to predict something. Yes, you can do a clustering and segmentation analysis to see, you know, who's close in embedding space and who, like, who might be like a patient that you think is the right kind of patient profile for a clinical trial. There's just a, you know, there's a lot of potential Downsord use cases for this kind of thing. That's right. Like, hey, look, here's a patient that was similar to you. Here is what happened with them, right? I think that could be a powerful thing. Yeah, very interesting. So you've done a lot of work on these models, but you've chosen to open source them. Could you give us an insight into why you decided to go to the open source route? Yeah, for now, we, honestly, it comes down to validation in benchmarking, right? So we need models like these to be broadly benchmarked by not just one academic medical institution, no matter how, how prestigious, not just 10, right? Like by the end of the year, we want this validated in one way, shape, or form on 100 different academic medical centers. I think that's absolutely doable, but I think it's also really important. It makes it kind of frustrating as a scientist. If I put my scientist, not my CEO had on, but I'm like, well, look, here's this other model. Is it better than ours? Is it not better than ours? Is there a place where there's an overlap? And we make a comparison, which one should I use? Well, I can't get their data and I can't get their models. So I don't know, right? The only thing we could do, you know, when face is like, you can't control what other people do, you can only control what you do. And what we can do, we can give them our model. And we could say, hey, look, at least you're aware of it. You can take it. You can benchmark it. And we will help you. We'll dedicate hours from some of that, you know, finance, dental engineers that exist in the world. And we'll help you do that. And that makes me feel pretty good at the end of the day. What I'm interested in from your perspective, so you already mentioned, you know, the question that you're sort of oncology forward, you know, how would you want people to think about, you know, sort of the limitations of the standard model bio models that are being released? Yeah. So right now, they're oncology forward. We chose that for one business reason. Right? Oncology is, you know, huge market. But it's also kind of the thin end of the wedge when it comes to precision medicine. You will have people write grants that say, as done in precision oncology, for instance. So we knew that was a good place to start if we wanted to work with, you know, people that have been thinking about this very deeply for many years. And that means that if you're using it and it should be oncology adjacent, it doesn't have to be oncology right now, right? Like in the sense of if you want to predict cardio toxicity, maybe that would make sense even though that's not necessarily a cardio thing. If you wanted to use your own cardio model to predict how things would happen in an oncology space that could work. We are moving into general chest disease with a collaborator that we'd love to announce soon, which we're really, really, really excited about. And I think that's going to be a pretty powerful thing because it's going to get us into some of the immune related things, which are also cancer relevant, right? The interstitial lung disease can be a devastating side effect to certain immunotherapies. We are thinking about oncology today in a couple months expected to be a lot more. Yeah. Yeah. So should I think about this as like you have a model architecture that you think is sort of universally applicable that models the patient, models the patient in a really good, in a really great way, but that your journey is a journey of incorporating more, you know, more patients with different therapeutic, therapeutic issues and potentially which weight differently toward the modalities that you're exposing the model toward based on, you know, what eat that cohort of patients and their journey through their patient experience. Is that, is that, is that, yeah? And if someone's like, look, I would love to use this, but your model's not working that well on it. Like we want to know. And sometimes we use that to identify what data to go after next, right? If someone's like, Hey, look, I really wanted to use this on, you know, predicting myocarditis and I can't, well, tell us and we can go try to pump that down. Well, Kevin, you've been an excellent guest. I really appreciate the time. If listeners want to learn more about your methods or, you know, the work that you're doing or where you're heading from an organizational perspective, where should they start? You know, they can, honestly, they can go to a substack. That's where you publish a lot of stuff, even before it makes it into papers. You know, so blog does standard model.bio. That's one of our favorite places or just email us. Yeah. Yeah, I highly recommend that substack. That was how I discovered your work and really enjoyed having you on to talk about it today. So thank you very much and look forward to connecting down the line. All right. Thank you, Ross. And that's it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate or leave a review in your podcast platform of choice. See you next time.

Podcast Summary

Key Points:

  1. Kevin Brown transitioned from pure math to brain-computer interfaces and deep learning, eventually joining Standard Bio to build multimodal foundation models for biology.
  2. A key "GPT-3 moment" came from realizing that scaling data and combining multiple modalities (e.g., imaging, genomics, EHR) could lead to powerful foundation models for patient trajectories.
  3. Standard Bio’s approach rejects treating a patient as a document; instead, it uses a JEPA-style model to map diverse modalities into a shared latent space, predicting future patient states rather than just next words.
  4. Each modality (e.g., radiology images, clinical notes, whole genome data) is processed by its own encoder, allowing flexible integration even when not all modalities are available for a given patient.
  5. The temporal nature of the model enables counterfactual reasoning, such as predicting how an intervention might alter a patient’s trajectory toward better outcomes.

Summary:

In this podcast, Kevin Brown, co-founder of Standard Bio, describes his journey from pure mathematics to building multimodal foundation models for biology. He highlights a pivotal "GPT-3 moment" when he realized that scaling data across modalities—such as neural recordings, imaging, and genomics—could transform drug development and patient care. , CT scans, pathology images, EHR notes, whole genome sequences) into a shared latent space.

Each modality is processed by its own encoder, allowing the model to handle missing modalities gracefully. The key innovation is temporal prediction: rather than predicting the next word, the model forecasts a patient’s future state in this abstract space, enabling counterfactual reasoning about interventions like treatments or lifestyle changes. , adverse events), offering a more holistic and dynamic understanding of health.

Brown emphasizes the importance of domain expertise and biostatistical rigor in building these models, aiming to create a foundation that facilitates any downstream task, from therapy selection to personalized care.

FAQs

Kevin Brown started in pure math, worked in a brain-computer interface lab, and joined Siemens Healthineers for computer-aided diagnosis, later moving to Bristol-Myers Squibb and then founding Standard Bio.

Reading the GPT-3 paper showed that scaling data and parameters could dramatically improve performance across many tasks, leading him to believe the same would happen in biology with multimodal foundation models.

He wanted to build a general foundation model using diverse data, but pharma data is biased toward their own drugs, limiting generalizability, so he founded Standard Bio to access broader data.

It means patients are complex, multimodal, and temporal, not just text notes. Their health involves images, genomics, and other data that can't be fully captured by language models alone.

Each modality (e.g., radiology, pathology, genomics) uses its own encoder to map data into a shared latent space, allowing the model to learn patient trajectories without requiring all modalities for every patient.

JEPA (Joint Embedding Predictive Architecture) predicts future patient embeddings in a shared space over time, enabling counterfactual reasoning and therapy effectiveness assessment by modeling patient trajectories.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.