Regina Barzilay, AI Small Molecule Discovery and Synthesis
28m 56s
Regina Barzali from MIT's Machine Learning for Pharmaceutical Discovery and Synthesis Consortium discusses the application of AI in drug discovery. The podcast highlights the importance of standards and benchmarks in AI research for drug discovery, emphasizing collaboration with industry for automating small molecule discovery. BenchSci's AI application in biomedical research is also promoted. Regina explains her transition from natural language processing to chemical synthesis and the unique perspective it offers. The MIT consortium aims to bridge the gap between academia and industry in drug discovery. There is a focus on the need to validate benchmark datasets in drug discovery research. The consortium provides public availability of research tools while offering additional support for consortium members to tailor tools to their specific needs.
Transcription
4405 Words, 25541 Characters
(upbeat music)
- Hello, and welcome to the Artificial Intelligence
in Drug Discovery podcast.
My name is Simon Smith, and I'm your host.
On this episode, I speak with Regina Barzali
of MIT's Machine Learning
for Pharmaceutical Discovery and Synthesis Consortium.
It's an honor to speak with Regina.
She's a prolific researcher
with many highly cited papers in machine learning,
particularly natural language processing,
oncology, and chemistry.
She's also a MacArthur Fellowship recipient.
In her role with the relatively new consortium,
she and her colleagues are collaborating with industry
to design software for automating
small molecule discovery and synthesis.
On this episode, you'll learn
what drove the need for the consortium,
the importance of standards and benchmarks
for AI drug discovery research,
and the progress researchers are making
in applying machine learning to chemical synthesis.
This episode is brought to you by BenchSci.
BenchSci uses artificial intelligence
to reduce the time, uncertainty, and cost
of biomedical research.
Use it to find research antibodies
up to 24 times faster than using PubMed or Google Scholar.
Just enter a protein of interest
and filter by technique, organism tissue,
or 15 other options.
BenchSci returns only relevant published figures and products.
Researchers in 14 of the top 20 pharmaceutical companies
in more than 1300 academic institutions
now rely on BenchSci to find antibodies.
It's free for researchers in academic
and nonprofit institutions.
You can sign up at BenchSci.com.
If you work in industry,
just use the contact form on BenchSci.com
to reach out for a demo.
And now on to the interview.
- Hi, Regina, welcome to the podcast.
- Hey.
- So thanks for joining me.
I wanted to start
by talking a bit about your background.
So as we were talking just before the call started here,
I went through looking at some of your papers
and you're one of the most prolific researchers
that I've seen, and particularly in natural language
processing with many highly cited papers
that long predate the hype cycle
that we're in now for machine learning.
But I wanted to know how you got from your work
on natural language processing to your current work
on small molecule discovery and synthesis.
Can you trace the path from that,
from the natural language processing
to the work you're doing on now with the MIT consortium?
- So in terms of the technique themselves,
there's actually a lot of connection
between natural language and design of small molecules.
And if you look at the field within machine learning,
natural language processing, which makes it unique,
is actually understanding the structures.
Natural language processing is primarily trees,
but it can be graphs.
And how do you translate the insides of linguistic theory
into the right structures
and then do learning and inference of those structures?
So in some ways or two,
when we're talking about molecules,
we're talking primarily about graphs,
but technically there is a lot of connection
between these two areas
because both operate over structural objects.
But my personal foray into this area was actually driven.
It was very random move because a few years ago,
faculty in chemical engineering approached me
and asked me if I would be interested to join him
at that proposal on using machine learning
for retro synthesis.
And my original role was to extract some information
from text.
I said, yes, of course, why not?
And then when we were afforded this proposal,
then it came to all the team members,
this understanding that it is really,
we can do much more than just extract from text.
We can start doing the prediction
on the molecules themselves.
And this was really kind of really exciting new area for me.
And from there, the things developed to the current state.
- And do you think that because of your background,
you have a unique perspective
that's gonna be different than somebody
who might be approaching this,
who has a background in chemistry
because you've worked with large document sets
in the way that you approach the problem?
Do you think that gives you a unique perspective on it?
- I think that unique perspective here,
and I really clearly see it in chemistry
is that the best models that you can create
actually utilize the particular properties
of the structures that you are working on
or some insight about the problem.
And if you look at the best works
that were developed overall
in the field of natural language processing,
so just taking something from machine learning
and applying it there.
But it's really understanding
what is unique about the structure,
how we can do inference more effectively
than in some general case.
And this type of skill,
you would translate it to a different structure
with a very different properties,
but that's where you can really
make interesting models which are highly effective.
And one thing that I noticed in this field
when I entered it,
there were very few papers on deep learning
and chemistry when we started here at MIT.
One thing that I noticed,
which was really not inspiring
is that people just took the model
that was developed for whatever,
and then you would like CNN and you just apply it.
And it's of course useful baseline,
and it tells us how difficult competition is a problem.
But I would say those are really not the most interesting
and promising models that are available out there.
And I wanna come back to talk in a bit
about what are the most promising approaches.
But before we go there,
what sparked me reaching out to you
was the announcement of the consortium.
How did that originate?
Where did the idea come from to bring together
both industry and researchers at MIT?
What sparked that idea?
- So we were doing a lot of work
on our DARPA project called Naked.
And it was great collaboration between professors
on the AI sector, me and Tomi Akala,
and our students in clubs, Jensen,
and others from chemical engineering.
And I absolutely have to mention Kona Kuli,
who was a student who really is in between machine learning
and chemical engineering,
even though he's a chemical engineering student.
So it was fun.
And then we kind of understood
that even though our original goal
was to do retro synthesis,
there are lots of other questions
that we're really interested in.
And if you ever visited Cambridge,
you can know it is at MIT, my own building.
I see Novartis from my windows.
And then we're surrounded by pharmaceutical companies.
And then I said, okay, all these questions
that we are already working on,
like property predictions, for instance,
are actual problems for that industry.
Why wouldn't we invite them over show what we have
and hear what kind of problems that we have?
And we had really fascinating conversation with them
that we understood, yeah, these tools can really change
how people are doing science and help them
in their drug discovery process.
And we learned a lot from this interaction.
And during this kind of, it was a workshop at MIT
saying last year, and last spring,
we realized it would be great to have consortium
because first of all, the questions that we're interested in
go way beyond the scope of regional DARPA call.
And second, in order for us
to really make innovations field,
it is not enough to run this model
on some standard benchmarks.
We need to understand what are the questions
and hear from them
and jointly to formulate agenda for this field.
And that's where we came up with an idea
of having consortium.
- Where there's specific things that you don't feel
or didn't feel were being addressed currently
by existing academic and industry initiatives or startups.
So for example, I know there's the Adam Project,
which is a number of companies like GSK
and Lawrence Livermore National Laboratory
have come together on that one.
And then there are at least 94 startups
that I've identified in this space
and of course a number of academic institutions.
Was there a big gap?
Was there some kind of gap that you felt existed
that the consortium really wanted to focus on?
- So let me just give you an idea
of how I feel there is a huge gap and opportunity.
If you look at the top machine learning conferences
like ICML and NEET,
papers on chemistry and dark discovery
are still a very tiny portion support for all papers.
Like you can't even compare
natural English processing in this field.
And this is tremendously important field for all of us.
You know, at some point we're all gonna be sick
and old and looking in the drives like.
So this is an important field.
And even you said you counted 94,
count how many startups you have in natural language processing.
(laughs)
Make any sense that in this field
you have so very few.
And again, I can count on my two hands,
the number of academic groups
that I know is really producing innovative research
in this field.
And what I mean innovative is not just a chemist
who took CNN and applied to their problem,
but really thinking about new models.
And it absolutely has to change.
And what I felt that it is not enough for us
to imagine what the industry may be interesting in.
We really need to go talk to them
and understand what are their problem specification.
On which questions,
if we give them the tool they're gonna use it
and which question they will not use it
and how they're using it.
And if you're looking at the literature,
I didn't get any of those insights
and letting me be very specific here.
So I'm wearing these names,
but there is this very nice toolkit,
a chemistry toolkit that come from Stanford
on property prediction.
And they connected 14,
I think 14 or 15 benchmarks dataset on property prediction
and they compared the variety of models.
And what we actually took this datasets,
we ran our own models, developed MIT,
we did better than any of the existing models.
But to me, the question was the following.
How does it actually work in industry?
Why they don't use this model?
Maybe find the MIT model just became available
like few months ago, but there were other models
because none of the people that they talked
actually use this model.
And it's not clear to us
whether this benchmark,
some of them very artificially created,
really a representative of the type of problems
that people in industry have.
We don't know that if there is a rank of the models
that we obtain by looking on the public benchmark,
how do they translate to the industry performance?
We don't know.
So the only way for us to know
is to put these tools in the industry
where they apply it in their own datasets.
And then here back,
it's actually what we're currently doing
with Amgen and Novartis and others,
where they take these tools
and run it on their selection of datasets,
and we will compile it in a hope to write a paper there,
which really will tell us,
are the standard datasets on which we are testing
predictive or real performance?
And without industry players,
we cannot answer these questions.
- So your work then partly now is to validate
the benchmark datasets,
'cause we've seen how in other areas,
like you have ImageNet for image classification
and it's very clear whether or not it's working
or not 'cause it's classifying images
or squad for question answering.
There are these things that the people rely on
so that when we see people,
there's been progress.
- Well, let's see, there are a lot of people in there
that will pick who are very upset with squad.
This is a new performance on squad.
It's not really representative on many other tasks,
which are even in question answering, you know,
you can be great on squad,
but there is reading combination dataset
which really doesn't translate for various reasons.
And it's not clear to the communities
that by optimizing on squad,
we really are improving our capacity to process the document.
This is an open question,
but in natural language processing,
whoever has doubts can create a new dataset
or take some other datasets and make this comparison.
But because most of the real datasets
for property predictions are proprietary datasets
within the pharmaceutical company,
unless they are the ones who are gonna try and tell us,
we would not know.
- So that, is that a big,
so that potentially the biggest barrier right now
is that we don't know the validity
of the benchmark datasets that we're using?
- It's just one small set,
but the problem is that you know that, you know,
like you can say people are working so much
on the property prediction task,
lots of different models.
And you say, why are they not used in industry
most of this model?
I mean, maybe there are companies which do use it,
but there are lots and lots of big companies
to do a lot of property testing
and they don't use this kind of model.
They use something very different
and they actually spend some time looking around the companies
to understand what's going on.
So I can tell you,
they don't use many of these models.
So the question is why?
And unless we bring them there,
understand the performance,
understand whether the benchmark is representative
we wouldn't know.
So this is an important, actually,
part of the consortium of understanding
this type of questions and understanding the benchmarking.
Let me give you another example.
We had a paper, I think it was ICML on lead optimization.
And the only corpus that is available
for lead optimization, the only benchmark,
is really toy baby benchmark
that human chemists can do very fast.
So we published it and we did well and whatever.
But does it really mean
that with this improvement in this performance,
we really are solving a complex,
realistic lead optimization task?
So that's why it's so important
to bring it to the users
who are gonna be actually employing these tools,
see what happens in reality and update our tools.
Because what I feel happens now
is people in pharma,
vast majority of chemists are really spectacular
and amazing, have a lot of intuition and expertise.
But maybe much less expertise within machine learning,
they may understand how the basic model works,
but they don't have what people in computer science
and machine learning can propose.
And you kind of take the pitch and divide it into two parts.
And in order to see the whole thing,
you need to put them together.
And this is maybe one of the main goals of our consortium
is to put the two audiences together
to create something new
and to move the state of the art forward.
- And it sounds like you have a good understanding
of what the barrier is
in terms of not having that industry validation.
When you talk to the industry partners,
did they have a similar recognition?
Like, do they have a similar perspective to you
that the problem is that we don't know
if this is applicable and that's why we aren't adopting it?
Did they agree with your perspective?
- I think that there are various reasons
and many of these people,
they used to use their certain toolkits
and they continue to use their certain toolkits,
they deliver the value.
This is how the flow works.
And part of the consortium should do a lot of education
to explain what is available and how the tools work.
But the barrier today,
it's not easy to cross a barrier.
And one with actually a cultural divide or knowledge divide
that most of these people who are users of these tools
are not aware about alternatives
or maybe they are aware about alternatives
but not sure if they're gonna deliver.
So part of it is just actually brings the tools
and I said, what do you need to make it work?
Those are the options, that's how you can optimize it.
And then observe what happens
and these things don't happen the way we hope.
That they will, one can update and happen.
It's like back and forth process.
- So I'm getting a sense of what consortium members
get both the research side and the industry side.
What ultimately it will be the output
and what will be available only to consortium members
and then will anything be available to the public
or to any researcher who wants to make use of it?
- So we are currently like the tools
that we already delivered to the consortium
our heterosynthesis tools and our property prediction tools.
And many of the tools that we've developed
like for instance our heterosynthesis tool
is publicly available and you need to sign to get it again
but it's publicly available tool.
The only issue is that we don't give you a copy of the code.
You can run it and try it and it's available online
but you would need to bring your molecules
and run it on ourselves.
We promised you that we delete them, we don't copy them
but for most of the consortium members
it was important to have their own proprietary copy
so that they can train it on their own material
rather than on some reacces or whatever materials
we use for training our tools
because then it can adapt much better
to their own chemistry.
So this is one example where we give them something
which is more unique to them.
For many papers that we produced
if you looked at my papers
we always release the code and make it publicly available
but what the extra benefit that consortium members get
is once they download this code we can help them
to integrate it within their pipeline
and now we kind of experience our first interaction
of getting their feedback and improving the tools for them
because you can give generic tool and the algorithm
but typically companies have their own unique needs.
- And I'm sure there's a lot of work once you get feedback.
- We emphasize that every single paper
that comes out of the MIT group
the code is publicly available
and I think it's extremely important
that at least from the public data sets
we can compare and reproduce
and whoever wants to improve their results
or to compare with us, they have an option to do so.
And we report everything for better or for worse
again on the publicly available data sets.
So in terms of the science
we don't compromise the science as part of the consortium.
- Great, I wanna switch gears a bit here
and talk about the research itself.
So I've read that one of the issues
with some of the current approaches
to automated molecular design is validity.
So they'll generate molecular structures that may be novel
but they're just totally invalid.
Is there something fundamentally wrong with approaches
people have taken to date?
So I just, I'll do the one comparison I have in my head
'cause I've played around with a lot of text generation
using recurrent neural networks.
And you get these strings that are like hilarious
but make absolutely no sense
'cause they don't understand the syntax or context
or logic and they have no common sense.
And that's the comparison I'm doing in my head.
Is there an approach that just,
maybe the way people are approaching it
is never gonna get us to a point of validity
unless we incorporate some sort of a higher level
or a higher order understanding
of how things come together or the properties.
Is there something that we're missing right now?
- So this is actually a great question.
And I want to connect it to what we discussed earlier
is about just taking tools that were developed
for something and plugging them in
and hoping for the best,
which is clearly does, it's not the right approach.
And for this reason,
when we were working on the lead optimization problem,
it became very clearly that if you kind of take
the standard approaches you translated to,
the molecular vector, you optimize it in hidden space
and you generate, which sounds like an interesting idea.
What you generate is not very valid.
And that's why I was sort of about representation,
which really would enable you very clearly control
for generating, for producing valid chemical structures.
And in this case, we actually encoded,
part of our encoding translated the molecule
into a tree of a bigger substructures,
which are based on some chemical theory,
how this translation is done.
And then you represent molecule both
at the level of this high level substructures
and also at the low level as a graph.
And when you're doing your generation,
you first generate the modified graph
where you can check for validity.
You have a language to check for validity
of your generated structure.
And then you kind of produce all the smaller,
fill in the smaller details.
And in that case, we were able with all this extra checkups
and the structure, we were able to generate 100%
correct molecules.
So I think the approach here to solving this dilemma
of generating invalid structures is really very carefully
think about the encoding and decoding and their design.
And in contrast to natural language processing,
one very appealing thing about chemistry
there are actually a lot of information that is available.
I mean, there are much more, I would say, kind of theory.
And it's much more defined than producing, in some ways,
valid natural language string.
And one of the big questions for me
in deep for us for the whole team, not only me,
that is how you can, in the most effective way,
can utilize chemical theory, chemical knowledge
to do a better job, to integrate them
nicely with deep learning models.
And I think this will be crucial for improving
the performance.
So from an architecture perspective,
it doesn't sound like that you think something
like a generative adversarial network is really going to cut
it because it's still not going to incorporate
some of the theoretical concepts.
So it's going to have to be some kind of different architecture
approach.
Do you--
The architectures that we had, for instance,
I think holds these bigger components.
You just kind of know what you're looking for.
So why would you let the model run wild?
You can give it feedback.
It doesn't mean that we have to encode the theory like even
else, but to take some insights from what we know
has to happen and directly put it into the model.
So I think there is a lot of interesting exploration
in the space of architectures and how maybe even in an automatic
way, we can translate these constraints
into fruitful architectures.
Does it limit the novelty, though,
when you do that level of constraint?
Or does it force you to incorporate certain biases,
maybe things that we've--
or rules or generalities that we've inferred that may be
artificially limiting?
Like, is there a trade-off between novelty and validity?
I think that we are not talking about over-constraining,
which is thinking about the way introducing
architectural bias.
The model can still push it into some new direction.
But if you know that something has to happen,
there is no point to make the model learn it,
because we know that this is true.
And it's particularly important in other areas
that you didn't mention.
But I know that this is a big case for the consortium members.
And for us, it's a case of transfer learning.
Because most of the time, if you're in a farmer,
you're actually focusing on new chemistry.
And our assumption that our training distribution fits
nicely-- this distribution doesn't really hold, correct?
So we've had many, many stories that people
are training it on some data, and they
apply it to really new chemistry.
And then the funny things that happen when you know how a classifier
predicts something with high confidence,
and it's totally off.
So in this case, incorporating some things
that we know in a smart way can be an additional way,
as a bias, not a hard-coded rule,
can help us to do better transfer.
Or I guess also it would help if everybody
was just willing to put all of their data
into a public space where you could just
build a more comprehensive model from many more data sets.
But I don't know how quickly that might happen.
Also notice we're bumping up against the time here, Regina.
So I've really enjoyed our talk.
And where can people go to learn more about you and your work?
So about me, you can just query, Regina Barzil, just one.
And you can read on the web page there
is a link to the work on chemistry.
And most of the information is available online.
And I also want to mention that I recorded pharma AI
tutorial, which is a 45-minute tutorial, which
kind of goes through the basics.
And I will put it on my web page for those
who want to learn in more detail what MIT team is doing.
Great.
I want to watch that myself.
So thank you so much for your time.
It took a while to coordinate this interview, I know.
But it was well worth it.
I learned a lot.
I'm sure our listeners did as well.
So thank you very much.
Thank you.
Thank you very much.
Bye-bye.
[MUSIC PLAYING]
You just listened to my conversation
with Regina Barzilai of MIT's Machine Learning
for Pharmaceutical Discovery and Synthesis Consortium.
I hope that you enjoyed it.
If you want to catch future episodes,
be sure to subscribe.
Just look for artificial intelligence and drug
discovery in your favorite podcast player.
Then hit the Subscribe button.
Until our next episode, be well and work smart.
[MUSIC PLAYING]
Podcast Summary
Key Points:
Regina Barzali of MIT discusses AI drug discovery research.
Importance of standards and benchmarks in AI drug discovery.
Collaboration with industry for automating small molecule discovery.
BenchSci's AI application in biomedical research.
Transition from natural language processing to chemical synthesis.
MIT consortium bridging gap between academia and industry in drug discovery.
Emphasis on validating benchmark datasets in drug discovery research.
Public availability of research tools from the consortium.
Summary:
Regina Barzali from MIT's Machine Learning for Pharmaceutical Discovery and Synthesis Consortium discusses the application of AI in drug discovery. The podcast highlights the importance of standards and benchmarks in AI research for drug discovery, emphasizing collaboration with industry for automating small molecule discovery. BenchSci's AI application in biomedical research is also promoted.
Regina explains her transition from natural language processing to chemical synthesis and the unique perspective it offers. The MIT consortium aims to bridge the gap between academia and industry in drug discovery. There is a focus on the need to validate benchmark datasets in drug discovery research.
The consortium provides public availability of research tools while offering additional support for consortium members to tailor tools to their specific needs.
FAQs
The need for the MIT consortium was driven by the desire to address questions that went beyond the scope of their DARPA project and to collaborate with industry to advance AI drug discovery research.
Industry validation of benchmark datasets is crucial to ensure that the tools developed are applicable to real-world problems and to understand how different models perform in industry settings.
Tools like the heterosynthesis tool and property prediction tools are publicly available, but consortium members receive proprietary copies to train on their own materials for better adaptation to their chemistry.
The MIT consortium ensures transparency by publicly releasing the code for every paper produced, allowing for comparison, reproduction, and improvement of results.
One challenge in automated molecular design is the generation of novel but invalid molecular structures. Approaches need to incorporate higher-level understanding to improve validity.
The main goal of the MIT consortium is to bring together academia and industry to create innovative solutions in AI drug discovery, addressing real-world problems and advancing the state of the art.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.