Go back

Dr. Luay Nakhleh: Modeling Evolution (Complexities, Cancer, The Role of Evolution in Bioinformatics)

25m 59s

Dr. Luay Nakhleh: Modeling Evolution (Complexities, Cancer, The Role of Evolution in Bioinformatics)

In the transcription, Dr. Lowei Nakhle discusses the importance of evolution in bioinformatics, emphasizing its role in providing a framework for analyzing biological data. The evolutionary lens allows for comparative analysis, revealing patterns in genomic data that are not visible through other methods. Understanding intricate biological processes like hybridization and horizontal gene transfer entails collaboration with biologists to develop accurate models. Dr. Nakhle highlights the iterative process of learning biology, breaking down complexity into manageable parts, and developing tools for specific challenges. His work in phylogenomics addresses incongruences in evolutionary patterns, while in computational cancer biology, he focuses on single-cell DNA data analysis. The approach involves learning the biology, refining statistical models, and adapting computational methods to address specific biological questions effectively.

Transcription

4286 Words, 24986 Characters

(upbeat music) - Welcome back to the Bioinformatics and Beyond podcast. We're joined once again by Dr. Lowei Nakhle, who's the JS Abercrombie Professor and Chair of the Computer Science Department. And we're gonna talk more about Dr. Nakhle's own work in Bioinformatics. Dr. Nakhle studies evolution. Thank you once again for being here. Let me just start with. And can you tell us a bit more about evolution? Why is it so central to Bioinformatics and to the work that you yourself have done? I mean, I know we talked about sequence alignment as sort of an introduction to Bioinformatics for a lot of, what does that have to do with evolution and why did you decide to start studying it? And tell us a little bit about your work. - Well, I think the best answer to your question is in a very famous quote that most people who are working in this area will use, inevitably they will be citing it or quoting it at some point, which is nothing in biology makes sense except in the light of evolution. And this is a quote from Dubjansky. So if you look at, I mean, evolution is the unifying theory of all of biology. And if you think about all the data that you'll look at or all the organisms you can see and not see and all the species around you, all the ecological systems, all of that is governed by this theory of evolution where all these species evolved from a common ancestor, a more social and common ancestor. And to me, if you wanna look at the biological question, maybe the best way to look at it is through this evolutionary lens. How would one go about looking at it through that lens? - So the way to look at it through that lens is usually in the comparative way. So when we think about evolution as a framework for answering questions or analyzing data, you wanna collect data from multiple individuals or from multiple species. And then evolution basically allows you to look at whatever question you are interested in, but now in a comparative way. So now you start looking, for example, at your question. So suppose you're asking a question about something related to the human genome. Well, if you now align the human genome with the chimp genome and with the gorilla genome and so on, so many patterns that if you look just at the human genome by itself, that you will not be able to see them. But once you start aligning them, so many things will start jumping at you without having to invoke fancy machine learning algorithms or anything like that. It's just because you're contrasting things just through that simple contrast, you can see certain patterns that exist. And again, without that evolutionary lens, you wouldn't have seen them. But the other perspective is that even when you think about these questions from a machine learning perspective or statistical inference perspective, putting evolution into the framework means that you are developing models that are evolution aware, that they are taken into account, that the data I'm looking at could have some temporal dependencies because they came from common ancestor and so on. So that allows you to start building these models that are generative by design so that they resemble the stochastic process that gave rise to the data. But they are also explainable, which is very important. And this is a big debate today in the area of machine learning and especially in deep learning about devising or developing explainable models. Explainable models are very important in biology, right? I mean, in biology, especially in these kind of genomics questions, it's usually not a classification problem. Is it category or category B? Usually we would like to understand what happened so that something now belongs to category A or to category B. So this is what I mean by taking an evolution perspective on things that the solution is evolution aware that, and again, it's not the cliche, it's when you are designing a model, a machine learning model, a statistical model, keep in mind that the data that you are looking at has evolved over time from a common ancestor. And you can start modeling that what happened to this data over time. Okay, mutations happen, mutations happen with this rate and so on. And to me, this is the most powerful framework to look at a biological question because sometimes one way to think about biological questions, say they are so complex, let's throw at them a very complex machine learning tool. Sometimes, a complex question, once you pause it in the right light, in the right lens, you can actually get the solution without having to throw all sorts of complex machine learning and math at the problem. And this is for me why it's very, very important to look at it from an evolutionary perspective. And evolution, in evolutionary biology, of course, that's all they do, but there are so many other fields now that take cancer genomics today. I mean, the issue of evolution and the tools from evolutionary biology and phylogenetics have made it into the field of computational cancer genomics now because they understand that when I look at these individual cells from a cancer patient, well, these cells in some sense, evolved from a common ancestral cell and there is a lot of benefit when you are trying to find the mutations and what happened to these cells in that patient. There's a lot of benefit to keep in mind that there is a tree, a tree structure underlying these cells that you just sequence because the cells divide and replicate and all of that. And they give rise to a tree structure. This is evolution at a small scale. This is evolution happening within the body of an individual. Again, we usually don't call it evolution. When we talk about Darwin's evolutionary theory, we're talking about species and so on. But what happened at the species level is what's happening at the cell level as well. So this is to me why it's so important to understand the theory of evolution and to try to apply it to any biological question you look at because today, biotechnology allow you to collect any data you want. Collecting data that will allow evolutionary analysis. I think it's very important. - In your own work, you've looked at a lot of the intricacies that come into play when you're looking at evolution and that are involved in the evolutionary process. Can you tell us about some of the different intricacies that arise in the evolutionary process? You mentioned, I think most people are aware of some of the basics. There's mutation that happens. There's inheritance where you pass on traits. There's, you might have a group of individuals, some species that kind of splits and diverges. What are some of the other intricacies and maybe feeding eventually into some of your own work that you've studied? - Yeah, so I think the easiest way to describe the main intricacies I focus on in my research is to think about that original sketch and in Darwin's notebooks. So when he actually was sketching evolution, he actually drew a tree and next to it he wrote, I think. So there's a sketch of a tree there. And I would say ever since that sketch, a lot of evolutionary biology, a lot of computational evolutionary biology and phylogenetics focused on the inference of a tree. As you said, when we infer a tree, we usually focus on the mutations, base pair, substitutions and so on that occurred on a genomic region as it evolved down that tree. Now, in the genomic or post-genomic era, today we have the ability to collect data on more than one gene or more than one genomic region. In fact, you can actually sequence the entire genome. But even without sequencing the entire genome, you can sequence multiple regions from across the genome. And some of the intricacies that have been highlighted recently and have been shown to be actually a very strong or very big challenge for inference and inferring evolution in this field is that when you look at different genomic regions, all of a sudden you start seeing that different genomic regions from the same species or the same set of genomes have different evolutionary histories, different trees. So to illustrate with an example, so if you look at the three species, the three primate species, human, chimp and gorilla, our understanding is that the way the species evolved is about four million years, roughly, for three to four million years ago, human and chimp split from a common ancestor. About eight to 10 million years ago, the ancestor of human and chimp split from the ancestor, split from gorilla, from the common ancestor. So if you think about that hypothesis, I can draw it as a simple tree that puts human and chimp next to each other as siblings in that tree, and then gorilla as their next relative. But if you start looking at genomic regions within the genomes of these three species, you will start seeing that some regions there actually give us the signal that chimp is closer to gorilla. Some other regions give us the signal that human is closer to gorilla. So now you are looking at a genome, one genome from human, one genome from chimp, one from gorilla, and you imagine that you are walking across the genomes, they are aligned and you are walking across them from left to right, so to speak. And as you are walking, you are looking at the tree that gave rise to that genomic region. So you come across a tree that puts human closer to chimp, then all of a sudden you jump into one that puts human closer to gorilla, then chimp closer to gorilla, and so on. So this is now what gave rise to a new field called phylogenomics, which is how do we infer evolutionary histories in the presence of this complexity? So in this case, what happened with human chimp gorilla is that the theory or the hypothesis is that a process known as incomplete lineage sorting explains what happened here. But my research actually goes beyond that and even goes beyond trees. I am very interested in processes that don't fit on a tree, and biology has several of those. So for example, in eukaryotes and organisms like humans, a primates, not human itself, but in primates, in plants, in fish, in frogs, and so many groups of species, hybridization happens. Okay, so hybridization is the mating between individuals from different species. Another process that happens in almost all sexual species is that the process of recombination. In bacteria, there is a process that's ubiquitous across almost the entire prokaryotic branch of the tree of life, which is known as horizontal gene transfer. These processes, recombination, hybridization, horizontal gene transfer, they are different biological processes. How horizontal gene transfer happens in bacteria is a different process from how hybridization happens in plants, and these are very different from how recombination happens. But even within the human population, if you think about structure to the human population, that there are subpopulations. This is not a tree structure, it's a clean tree structure that subpopulations do not interbreed with each other. So there's also that mixture or gene flow happening between these subpopulations. So when you start thinking about all these processes, now you are thinking beyond the tree, which is something we call phylogenetic network. So one of the big intricacies today is that when you look at the whole genome data or genome white data, and you're trying to understand how evolution happened, you need to account for this fact that different, that yes, there is one evolutionary history for how the species evolve, but no, there isn't a single evolutionary history of how all the genomic regions evolve. Different regions could have their own evolutionary histories. And now you need to start inferring these evolutionary histories of their species, and of the individual regions from the genomic data. And that's where most of my work focuses on in this domain, in the domain of phylogenomics, which is coming up with statistical models or probabilistic models that account for these kinds of processes, like incomplete lineage sorting I just mentioned, like hybridization and so on, and then developing methods for inference under such models. So this is one of the big fields that we are working on, and I have been working on this for some time now. Most recently in the last four or five years, I also got into the area of computational cancer biology, mainly focusing on evolutionary analysis of single cell DNA data. So this is relatively speaking a new technology that allows biologists to sequence the DNA of individual cells. And now we are interested in understanding the evolutionary history of the cells that have been sequenced because we wanna know how they split from each other, from their common ancestors, where the mutations happened, at what point, and so on. But there, the challenges are different. There, we don't have this phenomenon of different regions having different trees. They are most of the complexity or intricacies they have to do with the amount of signal in the data, which is usually weak, and the amount of noise in the data, which is very large, especially given that this technology we work with is relatively new and still has a high error rate. So these are some of the questions I work with. I also did work several years ago. I'm not, I don't have active projects in now, also looking at biological networks, like protein interaction networks, and so on, from an evolutionary perspective. And 'cause that's another powerful tool to me to try to find interesting patterns in biological interaction networks, is to look at networks across species, to try to find what parts have been conserved, and so on. - What's your process when you tackle all these different problems? This is a lot of different problems, and I feel like maybe following up from our previous conversation, there's a lot of modeling that goes into all of this. Is it kind of a, do you have sort of a standard procedure when you go into each of these ones? For instance, let's just start with the phylogenetics, inferring the evolution. You're gonna change it from a tree structure, purely branching, diverging to like a network structure where you take into account gene flow, horizontal gene transfer. How do you go about that process? - I mean, yes, I have a process, and the process always starts by learning the biology. So when you say I wanna model hybridization, you cannot just say I wanna model, I've heard of the word hybridization, let me model it. It doesn't work like that. You have to go speak to biologists, who are experts in hybridization. You have to read papers on hybridization. You have to read papers that talk about the process, how it happens, what are the consequences of the process, what happens after hybridization? Okay, when hybridization happens, okay, there are two individuals from different species meeting. Okay, how do the sibling, what type of genomic data they inherit from both parents? What happens to that genomic data after some period of time? The selection work on parts of it, how does recombination affect it? So there is a lot of time needs to be spent on understanding the biology, understanding the process. And this is one of the nice things about the field of informatics. You don't have to go now and become a biologist. It's a very collaborative field. It's a very interdisciplinary field. And now you can go to talk to biologists, who are the experts on that? And you start developing the model, and you go through a process of back and forth. You develop the model, you go talk to them, they say, no, this doesn't happen like this, no, you missed this, you missed that. And you keep refining the model. I mean, remember that there is this famous quote, and it is in the biology, it couldn't be more true, which is all models are wrong, right? All models are wrong, but some are useful. And in biology, every model we develop is wrong, for sure is wrong, right? Because biology is so complex, especially when you're looking at evolutionary biology, and trying to model processes that have been acting for millions or billions of years. So you have to balance these two things. One is try to learn the complexity of biology, but also at the end, you don't wanna keep waiting until you have the full picture, 'cause that will never happen. We will never, at least as I will not ever know, in my lifetime, the full picture of hybridization and all the intricacies of how, for example, evolution, plans evolve. So you have to strike the balance between learning enough so that you get down to the modeling and the inference to make progress. And then you make progress on that, and then all of a sudden, you know, some biologist is trying to use your tools, and they say, well, your tools don't work for my data. And you ask, why doesn't it work? For example, say, well, your tools assume what you, assume that the hybridization is deployed, that all species have two copies of the chromosome, including the hybrid. And the biologists say, but my data is polyploid. And as a computer scientist, the first question you ask, okay, what is polyploid, right? And this is where you get into this, what is polyploid? And then you can actually, once you get into that, basically all hell breaks loose because you can see the complexity, for example, in plans. It's not one polyploid, I mean, you can have four copies and six copies and eight copies and 14 copies of the chromosome. At that point, you start again thinking about, do I wanna develop for every type of polyploid, you know, that general, all-encompassing model, or do I wanna develop a tool that work for specific polyploys, let's say tetrapoleploids. You have to make these decisions to make progress, right? So it's this process of learning the biology. As you learn the biology and you say, okay, this is very complex, let's break it down into small chunks and let's develop tools for this. And as you develop this, to me, this is where I start learning computational work, right? And this is where if someone is prepared well with strong computational foundation, then when you are developing the solution, all of a sudden you find, okay, now I need to, you know, to rely on something from statistical inference. If you have strong foundation, you should be able to learn it, adapt it to your problem. And yeah, this is how I basically think about it. This is the process I follow. - Let's wrap up with this. So let's talk a little bit more about your specific work that you've been involved in. Let's talk about these specifics that you picked up that you decided you were gonna tackle some problem in these areas. So let's take, for instance, the horizontal gene transfer, hybridization, gene flow in evolution, and also when you looked into cancer evolution. For both of these, there were problems and issues and you learned about them and developed software and algorithms to address these. Can you tell us a little bit about what some of those were for both of those fields and what you ended up doing for those? - Yeah, so I mean, I can, for example, from the area of phylogenomics. Again, we started looking at data where the biologists said, I'm seeing all these incongruence patterns in my data. They don't seem to fit a tree pattern. So again, I don't wanna get into the details here, but you usually think about, okay, if it's incomplete lineage sorting, okay, I expect a certain distribution of gene trees and so on. Again, I don't wanna get technical into this. So they would look at the data and say it doesn't fit the distribution. It doesn't fit the expectation that I would have under a tree-like process. So now they start thinking, okay, maybe it is hybridization and this is what led me there, you know, the sense of, okay, I mean, there are all these biologists out there who have data. Their data doesn't, it didn't evolve down trees and we need to model it. And when we got into this, my first approach to the problem was a parsimony approach. Let's actually find the simplest possible solution. But then we started finding issues with the parsimony solution. For one thing, it can be arbitrarily far from the truth. For a second issue was that we couldn't estimate certain parameters that were of interest to the biology. This is a mathematical fact. We could not estimate parameters under them. It's not about the amount of data. It's just the parsimony framework doesn't allow us to estimate these parameters. Okay, so we said parsimony is underestimating things. Parsimony is not allowing us to do parameter estimation. How do we solve these things? We go to statistical approaches. So we went to statistical approaches and my first thoughts were maximum likelihood. Let's define likelihood function for the model. And let's find, you know, our inference will basically be, let's find the model that maximizes the likelihood. And then when I started working on that, and then after that I branched into Bayesian inference. Now, I always mentioned this story that I say that I didn't go from likelihood to Bayesian because that's just the standard path that you take. You first work on likelihood, then you develop Bayesian. No, I am not that type of person. I'm not interested in that type of research because Bayesian is there, let's do Bayesian. But I started asking questions about, you know, we have some biological knowledge about, for example, that if you take two species that have diverged for 400 million years, they cannot hybridize. I mean, biologists will tell you it makes no sense that they hybridize. Well, I start saying, okay, how do we incorporate this into our inference? And that's what led me to Bayesian inference because now I start, okay, through the prior, we can start putting this kind of knowledge that biologists have, or sometimes the biologists have knowledge about the hybrid itself. So you can start putting this information and that led me to Bayesian inference. So how I went from porcimony to likelihood to Bayesian, it was by necessity. It was not by, we ran out of computational problems on porcimony, let's invent a new field. No, porcimony was just too weak. We went to likelihood. Likelihood did not allow me to specify prior knowledge on that biologists know and they will tell you. And this is how we went to Bayesian. So this is what led me in that. In the area of cancer genomics, it's a close collaboration with colleagues at MD Anderson. And there, you start looking at the data and you basically, they tell you what the data, the complexity was not about the biology there, the complexity was about the data. So they start telling you, okay, we have the data. I said, okay, why don't we apply just standard phylogenetic methods? You have 200 genomes, run the phylogenetic method. Said, oh no, we cannot do that because there is error in the data, which is not something we usually deal with when we run neighbor joining or phylogenetic method. So that's what led us to, okay, let's now go into statistical modeling of error in the data and so on. So for me, the way I go into a certain path, it's driven by necessity, not because I wanna invent some fancy computational problem to keep me busy for a few years or to write a grant. Every problem I work on, I was convinced that I need to go to this specific type of solution to solve it. - All right, well, thank you very much. Once again, I really appreciate you taking the time to be here for this discussion today. Thank you, Dr. Nakhveh. - My pleasure, thank you so much, Leo. (upbeat music) - Thanks for tuning in to another episode. I hope you found it interesting. If you did, if you learned something new and you enjoyed the show, I'd love to hear about it on Twitter. You can join the conversation and keep up with the newest episodes and past guests by following @BioInfoPod. Feel free to tweet @theshow or send a DM about anything you liked, didn't like, who or what you'd like to see next, questions for future guests, or just chat about all things bioinformatics and of course, beyond. It really does make my data see people share on Twitter when they found the podcast useful. So definitely keep it coming. And again, that's @BioInfoPod. Finally, you can always help out by subscribing to the show, giving it a rating, or just recommending it to a friend who's interested in these topics. Thanks again and see you next time.

Podcast Summary

Key Points:

  1. Evolution is central to bioinformatics, providing a framework for analyzing data and answering questions.
  2. Evolutionary lens allows for comparative analysis, revealing patterns in genomic data.
  3. Understanding biological processes like hybridization and horizontal gene transfer requires collaboration with biologists.
  4. Modeling in bioinformatics involves learning the biology, developing statistical models, and balancing complexity and progress.

Summary:

In the transcription, Dr. Lowei Nakhle discusses the importance of evolution in bioinformatics, emphasizing its role in providing a framework for analyzing biological data. The evolutionary lens allows for comparative analysis, revealing patterns in genomic data that are not visible through other methods.

Understanding intricate biological processes like hybridization and horizontal gene transfer entails collaboration with biologists to develop accurate models. Dr. Nakhle highlights the iterative process of learning biology, breaking down complexity into manageable parts, and developing tools for specific challenges.

His work in phylogenomics addresses incongruences in evolutionary patterns, while in computational cancer biology, he focuses on single-cell DNA data analysis. The approach involves learning the biology, refining statistical models, and adapting computational methods to address specific biological questions effectively.

FAQs

Evolution is central to Bioinformatics as it provides a unifying theory for all of biology and helps in understanding patterns and relationships among different species.

Evolution allows for a comparative analysis of data from multiple individuals or species, revealing patterns that may not be apparent otherwise and aiding in the development of generative and explainable models.

In addition to mutation and inheritance, processes like incomplete lineage sorting, hybridization, recombination, and horizontal gene transfer introduce complexities in inferring evolutionary histories.

The process starts with learning the biology and collaborating with experts in the field to develop and refine models that balance biological complexity with computational feasibility.

In phylogenomics, addressing incongruence patterns in genomic data due to different evolutionary histories of regions. In cancer evolution, developing tools to analyze single-cell DNA data to understand evolutionary histories of cells with high noise levels.

The process involves understanding the biology of the problem, collaborating with domain experts, breaking down complexities into manageable chunks, and developing computational tools that balance biological realism with computational feasibility.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.