Dr. Luay Nakhleh: Modeling Evolution (Complexities, Cancer, The Role of Evolution in Bioinformatics)
25m 59s
In the transcription, Dr. Lowei Nakhle discusses the importance of evolution in bioinformatics, emphasizing its role in providing a framework for analyzing biological data. The evolutionary lens allows for comparative analysis, revealing patterns in genomic data that are not visible through other methods. Understanding intricate biological processes like hybridization and horizontal gene transfer entails collaboration with biologists to develop accurate models. Dr. Nakhle highlights the iterative process of learning biology, breaking down complexity into manageable parts, and developing tools for specific challenges. His work in phylogenomics addresses incongruences in evolutionary patterns, while in computational cancer biology, he focuses on single-cell DNA data analysis. The approach involves learning the biology, refining statistical models, and adapting computational methods to address specific biological questions effectively.
Transcription
4286 Words, 24986 Characters
(upbeat music)
- Welcome back to the Bioinformatics and Beyond podcast.
We're joined once again by Dr. Lowei Nakhle,
who's the JS Abercrombie Professor and Chair
of the Computer Science Department.
And we're gonna talk more about Dr. Nakhle's own work
in Bioinformatics.
Dr. Nakhle studies evolution.
Thank you once again for being here.
Let me just start with.
And can you tell us a bit more about evolution?
Why is it so central to Bioinformatics
and to the work that you yourself have done?
I mean, I know we talked about sequence alignment
as sort of an introduction to Bioinformatics for a lot of,
what does that have to do with evolution
and why did you decide to start studying it?
And tell us a little bit about your work.
- Well, I think the best answer to your question
is in a very famous quote that most people
who are working in this area will use,
inevitably they will be citing it
or quoting it at some point,
which is nothing in biology makes sense
except in the light of evolution.
And this is a quote from Dubjansky.
So if you look at, I mean, evolution is the unifying theory
of all of biology.
And if you think about all the data that you'll look at
or all the organisms you can see and not see
and all the species around you,
all the ecological systems,
all of that is governed by this theory of evolution
where all these species evolved from a common ancestor,
a more social and common ancestor.
And to me, if you wanna look at the biological question,
maybe the best way to look at it
is through this evolutionary lens.
How would one go about looking at it through that lens?
- So the way to look at it through that lens
is usually in the comparative way.
So when we think about evolution as a framework
for answering questions or analyzing data,
you wanna collect data from multiple individuals
or from multiple species.
And then evolution basically allows you to look
at whatever question you are interested in,
but now in a comparative way.
So now you start looking, for example,
at your question.
So suppose you're asking a question
about something related to the human genome.
Well, if you now align the human genome
with the chimp genome and with the gorilla genome
and so on, so many patterns
that if you look just at the human genome by itself,
that you will not be able to see them.
But once you start aligning them,
so many things will start jumping at you
without having to invoke fancy machine learning algorithms
or anything like that.
It's just because you're contrasting things
just through that simple contrast,
you can see certain patterns that exist.
And again, without that evolutionary lens,
you wouldn't have seen them.
But the other perspective is that even when you think
about these questions from a machine learning perspective
or statistical inference perspective,
putting evolution into the framework
means that you are developing models
that are evolution aware,
that they are taken into account,
that the data I'm looking at
could have some temporal dependencies
because they came from common ancestor and so on.
So that allows you to start building
these models that are generative by design
so that they resemble the stochastic process
that gave rise to the data.
But they are also explainable,
which is very important.
And this is a big debate today
in the area of machine learning
and especially in deep learning
about devising or developing explainable models.
Explainable models are very important in biology, right?
I mean, in biology,
especially in these kind of genomics questions,
it's usually not a classification problem.
Is it category or category B?
Usually we would like to understand what happened
so that something now belongs to category A
or to category B.
So this is what I mean by taking an evolution perspective
on things that the solution is evolution aware
that, and again, it's not the cliche,
it's when you are designing a model,
a machine learning model, a statistical model,
keep in mind that the data that you are looking at
has evolved over time from a common ancestor.
And you can start modeling
that what happened to this data over time.
Okay, mutations happen, mutations happen
with this rate and so on.
And to me, this is the most powerful framework
to look at a biological question
because sometimes one way to think
about biological questions, say they are so complex,
let's throw at them a very complex machine learning tool.
Sometimes, a complex question,
once you pause it in the right light, in the right lens,
you can actually get the solution
without having to throw all sorts
of complex machine learning and math at the problem.
And this is for me why it's very, very important
to look at it from an evolutionary perspective.
And evolution, in evolutionary biology, of course,
that's all they do, but there are so many other fields now
that take cancer genomics today.
I mean, the issue of evolution and the tools
from evolutionary biology and phylogenetics
have made it into the field
of computational cancer genomics now
because they understand that when I look
at these individual cells from a cancer patient,
well, these cells in some sense,
evolved from a common ancestral cell
and there is a lot of benefit
when you are trying to find the mutations
and what happened to these cells in that patient.
There's a lot of benefit to keep in mind
that there is a tree, a tree structure
underlying these cells that you just sequence
because the cells divide and replicate and all of that.
And they give rise to a tree structure.
This is evolution at a small scale.
This is evolution happening within the body of an individual.
Again, we usually don't call it evolution.
When we talk about Darwin's evolutionary theory,
we're talking about species and so on.
But what happened at the species level
is what's happening at the cell level as well.
So this is to me why it's so important
to understand the theory of evolution
and to try to apply it to any biological question
you look at because today,
biotechnology allow you to collect any data you want.
Collecting data that will allow evolutionary analysis.
I think it's very important.
- In your own work, you've looked at a lot of the intricacies
that come into play when you're looking at evolution
and that are involved in the evolutionary process.
Can you tell us about some of the different intricacies
that arise in the evolutionary process?
You mentioned, I think most people are aware
of some of the basics.
There's mutation that happens.
There's inheritance where you pass on traits.
There's, you might have a group of individuals,
some species that kind of splits and diverges.
What are some of the other intricacies
and maybe feeding eventually into some of your own work
that you've studied?
- Yeah, so I think the easiest way
to describe the main intricacies I focus on in my research
is to think about that original sketch
and in Darwin's notebooks.
So when he actually was sketching evolution,
he actually drew a tree and next to it he wrote, I think.
So there's a sketch of a tree there.
And I would say ever since that sketch,
a lot of evolutionary biology,
a lot of computational evolutionary biology
and phylogenetics focused on the inference of a tree.
As you said, when we infer a tree,
we usually focus on the mutations,
base pair, substitutions and so on
that occurred on a genomic region
as it evolved down that tree.
Now, in the genomic or post-genomic era,
today we have the ability to collect data
on more than one gene or more than one genomic region.
In fact, you can actually sequence the entire genome.
But even without sequencing the entire genome,
you can sequence multiple regions from across the genome.
And some of the intricacies that have been highlighted recently
and have been shown to be actually a very strong
or very big challenge for inference
and inferring evolution in this field
is that when you look at different genomic regions,
all of a sudden you start seeing
that different genomic regions from the same species
or the same set of genomes
have different evolutionary histories, different trees.
So to illustrate with an example,
so if you look at the three species,
the three primate species, human, chimp and gorilla,
our understanding is that the way the species evolved
is about four million years, roughly,
for three to four million years ago,
human and chimp split from a common ancestor.
About eight to 10 million years ago,
the ancestor of human and chimp split from the ancestor,
split from gorilla, from the common ancestor.
So if you think about that hypothesis,
I can draw it as a simple tree
that puts human and chimp next to each other
as siblings in that tree,
and then gorilla as their next relative.
But if you start looking at genomic regions
within the genomes of these three species,
you will start seeing that some regions there
actually give us the signal that chimp is closer to gorilla.
Some other regions give us the signal
that human is closer to gorilla.
So now you are looking at a genome,
one genome from human, one genome from chimp,
one from gorilla,
and you imagine that you are walking across the genomes,
they are aligned and you are walking across them
from left to right, so to speak.
And as you are walking, you are looking at the tree
that gave rise to that genomic region.
So you come across a tree that puts human closer to chimp,
then all of a sudden you jump into one
that puts human closer to gorilla,
then chimp closer to gorilla, and so on.
So this is now what gave rise to a new field
called phylogenomics,
which is how do we infer evolutionary histories
in the presence of this complexity?
So in this case, what happened with human chimp gorilla
is that the theory or the hypothesis
is that a process known as incomplete lineage sorting
explains what happened here.
But my research actually goes beyond that
and even goes beyond trees.
I am very interested in processes that don't fit on a tree,
and biology has several of those.
So for example, in eukaryotes and organisms like humans,
a primates, not human itself,
but in primates, in plants, in fish, in frogs,
and so many groups of species, hybridization happens.
Okay, so hybridization is the mating
between individuals from different species.
Another process that happens in almost all sexual species
is that the process of recombination.
In bacteria, there is a process that's ubiquitous
across almost the entire prokaryotic branch
of the tree of life,
which is known as horizontal gene transfer.
These processes, recombination, hybridization,
horizontal gene transfer,
they are different biological processes.
How horizontal gene transfer happens in bacteria
is a different process from how hybridization happens
in plants, and these are very different
from how recombination happens.
But even within the human population,
if you think about structure to the human population,
that there are subpopulations.
This is not a tree structure, it's a clean tree structure
that subpopulations do not interbreed with each other.
So there's also that mixture or gene flow happening
between these subpopulations.
So when you start thinking about all these processes,
now you are thinking beyond the tree,
which is something we call phylogenetic network.
So one of the big intricacies today
is that when you look at the whole genome data
or genome white data,
and you're trying to understand how evolution happened,
you need to account for this fact that different,
that yes, there is one evolutionary history
for how the species evolve,
but no, there isn't a single evolutionary history
of how all the genomic regions evolve.
Different regions could have their own evolutionary histories.
And now you need to start inferring
these evolutionary histories of their species,
and of the individual regions from the genomic data.
And that's where most of my work focuses on
in this domain, in the domain of phylogenomics,
which is coming up with statistical models
or probabilistic models that account
for these kinds of processes,
like incomplete lineage sorting I just mentioned,
like hybridization and so on,
and then developing methods for inference
under such models.
So this is one of the big fields that we are working on,
and I have been working on this for some time now.
Most recently in the last four or five years,
I also got into the area of computational cancer biology,
mainly focusing on evolutionary analysis
of single cell DNA data.
So this is relatively speaking a new technology
that allows biologists to sequence the DNA
of individual cells.
And now we are interested in understanding
the evolutionary history of the cells that have been sequenced
because we wanna know how they split from each other,
from their common ancestors,
where the mutations happened, at what point, and so on.
But there, the challenges are different.
There, we don't have this phenomenon
of different regions having different trees.
They are most of the complexity or intricacies
they have to do with the amount of signal in the data,
which is usually weak, and the amount of noise
in the data, which is very large,
especially given that this technology we work with
is relatively new and still has a high error rate.
So these are some of the questions I work with.
I also did work several years ago.
I'm not, I don't have active projects in now,
also looking at biological networks,
like protein interaction networks, and so on,
from an evolutionary perspective.
And 'cause that's another powerful tool to me
to try to find interesting patterns
in biological interaction networks,
is to look at networks across species,
to try to find what parts have been conserved, and so on.
- What's your process when you tackle
all these different problems?
This is a lot of different problems,
and I feel like maybe following up
from our previous conversation,
there's a lot of modeling that goes into all of this.
Is it kind of a, do you have sort of a standard procedure
when you go into each of these ones?
For instance, let's just start with the phylogenetics,
inferring the evolution.
You're gonna change it from a tree structure,
purely branching, diverging to like a network structure
where you take into account gene flow,
horizontal gene transfer.
How do you go about that process?
- I mean, yes, I have a process,
and the process always starts by learning the biology.
So when you say I wanna model hybridization,
you cannot just say I wanna model,
I've heard of the word hybridization, let me model it.
It doesn't work like that.
You have to go speak to biologists,
who are experts in hybridization.
You have to read papers on hybridization.
You have to read papers that talk about the process,
how it happens, what are the consequences of the process,
what happens after hybridization?
Okay, when hybridization happens,
okay, there are two individuals
from different species meeting.
Okay, how do the sibling,
what type of genomic data they inherit from both parents?
What happens to that genomic data
after some period of time?
The selection work on parts of it,
how does recombination affect it?
So there is a lot of time needs to be spent
on understanding the biology,
understanding the process.
And this is one of the nice things
about the field of informatics.
You don't have to go now and become a biologist.
It's a very collaborative field.
It's a very interdisciplinary field.
And now you can go to talk to biologists,
who are the experts on that?
And you start developing the model,
and you go through a process of back and forth.
You develop the model, you go talk to them,
they say, no, this doesn't happen like this,
no, you missed this, you missed that.
And you keep refining the model.
I mean, remember that there is this famous quote,
and it is in the biology, it couldn't be more true,
which is all models are wrong, right?
All models are wrong, but some are useful.
And in biology, every model we develop is wrong,
for sure is wrong, right?
Because biology is so complex,
especially when you're looking at evolutionary biology,
and trying to model processes that have been acting
for millions or billions of years.
So you have to balance these two things.
One is try to learn the complexity of biology,
but also at the end, you don't wanna keep waiting
until you have the full picture,
'cause that will never happen.
We will never, at least as I will not ever know,
in my lifetime, the full picture of hybridization
and all the intricacies of how, for example,
evolution, plans evolve.
So you have to strike the balance between learning enough
so that you get down to the modeling and the inference
to make progress.
And then you make progress on that,
and then all of a sudden, you know,
some biologist is trying to use your tools,
and they say, well, your tools don't work for my data.
And you ask, why doesn't it work?
For example, say, well, your tools assume what you,
assume that the hybridization is deployed,
that all species have two copies of the chromosome,
including the hybrid.
And the biologists say, but my data is polyploid.
And as a computer scientist, the first question you ask,
okay, what is polyploid, right?
And this is where you get into this, what is polyploid?
And then you can actually, once you get into that,
basically all hell breaks loose
because you can see the complexity, for example, in plans.
It's not one polyploid, I mean, you can have four copies
and six copies and eight copies
and 14 copies of the chromosome.
At that point, you start again thinking about,
do I wanna develop for every type of polyploid,
you know, that general, all-encompassing model,
or do I wanna develop a tool that work
for specific polyploys, let's say tetrapoleploids.
You have to make these decisions to make progress, right?
So it's this process of learning the biology.
As you learn the biology and you say, okay,
this is very complex, let's break it down into small chunks
and let's develop tools for this.
And as you develop this, to me,
this is where I start learning computational work, right?
And this is where if someone is prepared well
with strong computational foundation,
then when you are developing the solution,
all of a sudden you find, okay, now I need to, you know,
to rely on something from statistical inference.
If you have strong foundation,
you should be able to learn it, adapt it to your problem.
And yeah, this is how I basically think about it.
This is the process I follow.
- Let's wrap up with this.
So let's talk a little bit more about your specific work
that you've been involved in.
Let's talk about these specifics that you picked up
that you decided you were gonna tackle some problem
in these areas.
So let's take, for instance,
the horizontal gene transfer, hybridization,
gene flow in evolution,
and also when you looked into cancer evolution.
For both of these, there were problems and issues
and you learned about them and developed software
and algorithms to address these.
Can you tell us a little bit about what some of those were
for both of those fields
and what you ended up doing for those?
- Yeah, so I mean, I can, for example,
from the area of phylogenomics.
Again, we started looking at data where the biologists said,
I'm seeing all these incongruence patterns in my data.
They don't seem to fit a tree pattern.
So again, I don't wanna get into the details here,
but you usually think about,
okay, if it's incomplete lineage sorting,
okay, I expect a certain distribution of gene trees
and so on.
Again, I don't wanna get technical into this.
So they would look at the data
and say it doesn't fit the distribution.
It doesn't fit the expectation that I would have
under a tree-like process.
So now they start thinking, okay, maybe it is hybridization
and this is what led me there, you know, the sense of,
okay, I mean, there are all these biologists out there
who have data.
Their data doesn't, it didn't evolve down trees
and we need to model it.
And when we got into this,
my first approach to the problem was a parsimony approach.
Let's actually find the simplest possible solution.
But then we started finding issues
with the parsimony solution.
For one thing, it can be arbitrarily far from the truth.
For a second issue was that we couldn't estimate
certain parameters that were of interest to the biology.
This is a mathematical fact.
We could not estimate parameters under them.
It's not about the amount of data.
It's just the parsimony framework
doesn't allow us to estimate these parameters.
Okay, so we said parsimony is underestimating things.
Parsimony is not allowing us to do parameter estimation.
How do we solve these things?
We go to statistical approaches.
So we went to statistical approaches
and my first thoughts were maximum likelihood.
Let's define likelihood function for the model.
And let's find, you know,
our inference will basically be,
let's find the model that maximizes the likelihood.
And then when I started working on that,
and then after that I branched into Bayesian inference.
Now, I always mentioned this story
that I say that I didn't go from likelihood to Bayesian
because that's just the standard path that you take.
You first work on likelihood, then you develop Bayesian.
No, I am not that type of person.
I'm not interested in that type of research
because Bayesian is there, let's do Bayesian.
But I started asking questions about,
you know, we have some biological knowledge about,
for example, that if you take two species
that have diverged for 400 million years,
they cannot hybridize.
I mean, biologists will tell you it makes no sense
that they hybridize.
Well, I start saying, okay,
how do we incorporate this into our inference?
And that's what led me to Bayesian inference
because now I start, okay, through the prior,
we can start putting this kind of knowledge
that biologists have,
or sometimes the biologists have knowledge
about the hybrid itself.
So you can start putting this information
and that led me to Bayesian inference.
So how I went from porcimony to likelihood to Bayesian,
it was by necessity.
It was not by, we ran out of computational problems
on porcimony, let's invent a new field.
No, porcimony was just too weak.
We went to likelihood.
Likelihood did not allow me to specify prior knowledge
on that biologists know and they will tell you.
And this is how we went to Bayesian.
So this is what led me in that.
In the area of cancer genomics,
it's a close collaboration with colleagues
at MD Anderson.
And there, you start looking at the data
and you basically, they tell you what the data,
the complexity was not about the biology there,
the complexity was about the data.
So they start telling you, okay, we have the data.
I said, okay, why don't we apply
just standard phylogenetic methods?
You have 200 genomes, run the phylogenetic method.
Said, oh no, we cannot do that
because there is error in the data,
which is not something we usually deal with
when we run neighbor joining or phylogenetic method.
So that's what led us to, okay,
let's now go into statistical modeling of error
in the data and so on.
So for me, the way I go into a certain path,
it's driven by necessity,
not because I wanna invent some fancy computational problem
to keep me busy for a few years or to write a grant.
Every problem I work on, I was convinced
that I need to go to this specific type of solution
to solve it.
- All right, well, thank you very much.
Once again, I really appreciate you taking the time
to be here for this discussion today.
Thank you, Dr. Nakhveh.
- My pleasure, thank you so much, Leo.
(upbeat music)
- Thanks for tuning in to another episode.
I hope you found it interesting.
If you did, if you learned something new
and you enjoyed the show,
I'd love to hear about it on Twitter.
You can join the conversation and keep up
with the newest episodes and past guests
by following @BioInfoPod.
Feel free to tweet @theshow or send a DM
about anything you liked, didn't like,
who or what you'd like to see next,
questions for future guests,
or just chat about all things bioinformatics
and of course, beyond.
It really does make my data see people share on Twitter
when they found the podcast useful.
So definitely keep it coming.
And again, that's @BioInfoPod.
Finally, you can always help out
by subscribing to the show, giving it a rating,
or just recommending it to a friend
who's interested in these topics.
Thanks again and see you next time.
Podcast Summary
Key Points:
Evolution is central to bioinformatics, providing a framework for analyzing data and answering questions.
Evolutionary lens allows for comparative analysis, revealing patterns in genomic data.
Understanding biological processes like hybridization and horizontal gene transfer requires collaboration with biologists.
Modeling in bioinformatics involves learning the biology, developing statistical models, and balancing complexity and progress.
Summary:
In the transcription, Dr. Lowei Nakhle discusses the importance of evolution in bioinformatics, emphasizing its role in providing a framework for analyzing biological data. The evolutionary lens allows for comparative analysis, revealing patterns in genomic data that are not visible through other methods.
Understanding intricate biological processes like hybridization and horizontal gene transfer entails collaboration with biologists to develop accurate models. Dr. Nakhle highlights the iterative process of learning biology, breaking down complexity into manageable parts, and developing tools for specific challenges.
His work in phylogenomics addresses incongruences in evolutionary patterns, while in computational cancer biology, he focuses on single-cell DNA data analysis. The approach involves learning the biology, refining statistical models, and adapting computational methods to address specific biological questions effectively.
FAQs
Evolution is central to Bioinformatics as it provides a unifying theory for all of biology and helps in understanding patterns and relationships among different species.
Evolution allows for a comparative analysis of data from multiple individuals or species, revealing patterns that may not be apparent otherwise and aiding in the development of generative and explainable models.
In addition to mutation and inheritance, processes like incomplete lineage sorting, hybridization, recombination, and horizontal gene transfer introduce complexities in inferring evolutionary histories.
The process starts with learning the biology and collaborating with experts in the field to develop and refine models that balance biological complexity with computational feasibility.
In phylogenomics, addressing incongruence patterns in genomic data due to different evolutionary histories of regions. In cancer evolution, developing tools to analyze single-cell DNA data to understand evolutionary histories of cells with high noise levels.
The process involves understanding the biology of the problem, collaborating with domain experts, breaking down complexities into manageable chunks, and developing computational tools that balance biological realism with computational feasibility.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.