Drug Discovery's AI Paradigm Shift with Dr. Abraham Heifets (Atomwise)
42m 16s
The AI Health podcast delves into the transformative impact of AI on healthcare, biotech, and medicine, with a focus on drug discovery. The episode discusses the escalating costs and timeframes involved in bringing new drugs to market, emphasizing the necessity for more efficient methods. It explains the drug discovery process, from target identification to molecule design, highlighting the role of machine learning in hit discovery. Atomwise's pioneering use of deep convolutional neural networks in drug discovery is explored, along with its collaboration with various industry players. The AtomNet technology is unveiled, showcasing its innovative approach using three-dimensional grids and color channels for biochemistry applications. The podcast sheds light on the significance of predictive models in optimizing drug development and the need for prospective testing to validate hypotheses effectively.
Transcription
7147 Words, 41105 Characters
This is the AI Health podcast, where we will explore the ways in which AI will transform
healthcare, biotech, and medicine through conversations with entrepreneurs, investors,
and scientists.
I'm your co-host Prana Vrajpurkar.
And I'm Adriel Seporta.
Welcome to our very first episode of the AI Health podcast.
Adriel, we'll be kicking off the podcast with our first episode on how artificial intelligence
is changing drug discovery.
I can't wait.
Before we get into our interview today with our guest, I want to share a little bit about
how the drug discovery process works and dive into an application of artificial intelligence
in the space.
I've heard it's very expensive and very time consuming.
Yes.
Do you want to take a guess as to how expensive it is to get a new drug to market?
Oh, man.
My guess is that it's in the billions.
That's right.
In recent years, the average price tag for getting a new drug to market has risen to
about $2.6 billion with an estimated delivery date of 10 to 15 years.
Wow.
One stat is that 10 years ago, every dollar invested in research and development saw a
return of $0.10.
Today it yields less than $0.02.
So everyone agrees that we need to bring down the time and the cost required for the development
of drugs.
Yeah.
That makes sense.
And then, of course, at the same time as costs are going up, we're seeing more and more
urgent global health challenges like emerging pandemic viruses or increasing antibiotic
resistance.
And it sounds like researchers need to find a way to shorten time to discovery and reduce
costs.
So tell me how this discovery process works.
Well, modern day drug discovery is a long and complex process.
Historically, pharmaceutical research started from, often through luck, observed medical
effects.
This paradigm of rational drug design turned this around by analyzing biological pathways
and identifying drugable targets.
So in the life sciences, when a protein is known to play an important role in a disease,
we call it a target.
Examples of drug targets include proteins that help tumors grow or proteins that viruses
use to infect human cells.
Now, once this target is identified, we have the problem of designing a drug that can safely
and effectively work on the target.
The goal is to come up with a molecule that can chemically bind or attach to the target
protein and modify it so that it no longer contributes to the disease or its symptoms.
These are usually small molecule drugs, which are typically composed of only 20 to 100 atoms.
And these tend to work well within the body, within cells.
There are other classes of medicine that exist and are being developed, but small molecules
are most popular, especially for computational approaches.
Got it.
So the first stage of drug discovery is all about finding the right target and then finding
the right molecule or drug that works on that target.
And then the later stages of drug development focus on finding how that drug first works
on animals, and then later on humans when the drug is administered to patients in a
clinical trial.
So what is machine learning doing for this?
So there are efforts throughout the drug discovery process that machine learning is helping with.
Our future episodes will cover other parts of the drug discovery process and, of course,
also topics outside of drug discovery.
But for today, I want to focus particularly on hit discovery.
So hit discovery helps in the identification of molecules with activity against the target.
What this means is whether a given molecule will bind to a target, and if so, how strongly.
And a good metaphor here is how keys fit into locks.
There are billions of possible keys, but only a few that open each specific lock.
And we're trying to find one of these.
Now it's important to remember that molecules that are good candidates as drugs still have
to jump through other hoops.
They have to make it through the gut into the bloodstream without immediately being broken
down.
They have to work in an organ without disrupting others.
They have to avoid binding to proteins that they're not meant to bind to.
And so there are other considerations beyond how well they bind.
Now to come back to determining whether a molecule will bind to a target, drug companies
have relied on conventional tools for this, with chemical analyses in test tubes to identify
molecules that interact with the drug target.
And there have been technological advances that enable what's called high throughput
screening, which includes equipment to rapidly test thousands to millions of samples for
biological activity.
And millions of compounds sounds like a lot, but this is actually a tiny, tiny fraction
of the chemical universe of drug-like synthesizable compounds, which is estimated to be around
10 raised to 60.
Wow.
10 to the 60 is a lot.
So is it possible to use computation to be able to select which compounds we should be
testing?
Yes.
The idea of using computation to what's called virtually screen compounds actually got a
big boost in the '70s and the '80s.
The idea here is to be able to use software to select compounds based on what's known
about experiments that worked or knowledge of the structure of the target.
Unfortunately, though, these early computational methods failed to live up to the hype, and
the field of computational drug discovery went through a sort of AI winter.
Funding and enthusiasm for computational methods have recently returned with the advent
of deep learning methods.
So what can these methods do?
Now with virtual screening, we can automatically evaluate very large what are called libraries
of compounds using computer programs.
So we can filter from the incredibly enormous chemical space of 10 raised to 60 conceivable
compounds to find a manageable number that can now be synthesized, purchased, and tested
in a lab.
And today, we're going to be interviewing Dr. Abraham Heifetz, the CEO and co-founder
of Atomwise, a company that has been at the forefront of using deep convolutional neural
networks for drug discovery.
So to give a brief background on Atomwise, Atomwise uses AI to computationally screen
over 16 billion molecules in less than two days.
Those 16 billion molecules is 5,000 times larger than typical big pharma corporate collections.
So we'll talk to Abe about the progress that Atomwise has been able to make on discovering
novel molecules today.
Abe, I'd love to start by asking you, what is the greatest challenge today in preclinical
drug discovery and development?
Sure.
Do you want the list alphabetical or chronological?
There's a lot of them.
I think people like to talk about preclinical drug discovery as a multi-objective optimization
problem.
That's a phrase that you'll hear.
What it means is that a drug has to have success on a number of factors.
It's got to stick to the disease protein that you want to shut down.
It's got to bounce off the proteins in your liver and your kidneys and your heart and
your brain that should keep functioning.
It's got to do that while being able to dissolve in water, dissolve in the gut, get through
the gut into the blood, stay soluble in the blood.
Blood has proteins that are breaking down drugs all the time, and so it's got to last
long enough in the body to get to the disease, and then it's got to get inside the disease
cells and operate on the cells.
There's a huge range of factors that a new medicine has to be able to pass, let alone
that when we start testing it, every drug that we make is actually a multi-species drug
because every drug that gets through FDA approval, it had to work in mice, and then it had to
work in rats, and then it had to work in large animals, and then it had to work in primates
usually, and only then did it have to work in humans, and so you actually have to design
something that hits all those factors for multiple different species.
I'd love to understand, when you started AtomWise, what was the space like that made you think
this is what the space needs in order for us to be able to do something different and
something that's going to make a positive impact?
Sure, so we started as a research project, so my co-founder and I, Ishar Valik, we were
students in the same laboratory at the University of Toronto, and we had the good fortune to
be there when modern machine learning was really being invented, so the computational
biology group, and we were students in the same lab, was on the same hallway as Jeff
Hinton's machine learning group, so we got to see in these hallway conversations around
sharing the same coffee pot, we got to see the success that convolutional neural networks
were delivering in image recognition and speech recognition.
Convolutional neural networks are still our species best technology today for those problems.
Your listeners have, if they've ever used Siri or Alexa, if they've ever uploaded a
photo to Facebook and said, Facebook said, "Hey, did you mean to tag your friend?"
If they've ever seen a self-driving car on the road, they've interacted with these convolutional
neural networks, and so we had the good fortune to see that kind of result early, and so we
were the first group to say what works on Jeff Hinton's side of the hallway, could work
on our side of the hallway and to apply speech recognition and image recognition to molecular
recognition.
I'm curious, so you mentioned Siri and Alexa as examples.
How long is it going to be before the drugs that we ingest are going to be products of
AI?
Sometimes I get an even more aggressive form of this question.
People say, "Look, I'll believe AI when I see the AI discovered drug.
Where's the AI discovered drug?"
You know what's funny is when we talk about AI today, probably we're talking about machine
learning, and when we talk about machine learning today, probably we're talking about statistical
approaches.
If you think about statistical approaches in drug discovery, and you say, "Where has
statistics ever discovered a drug?"
You could probably either believe that statistics has discovered no drugs, because you need
the biologists, and you need the chemists, and you need the clinicians, and you need
the doctors, and you need the patients.
It's also reasonable to say statistics has discovered every drug, because I dare you
to try to do any part of that without a deep and nuanced appreciation of statistics.
Try to do basic research, try to do a clinical trial.
You need statistics in every step, and I think that's why you see AI today being applied
to so many different pieces in drug discovery, baked in from the very beginning that design
was a fundamentally statistical design.
The machine on which these rules run is not a neural network on a GPU in a cloud, but
a neural network in the meat inside your skull, or your friendly medicinal chemist skull.
We can augment human capabilities and give humans superpower.
You can have the Iron Man suit, but for the brain, that's the way to use the neural network.
I love the idea of giving medicinal chemists superpowers.
What an amazing image.
My very layman's understanding of drug development is I think of these big, big pharma companies
sort of taking things end-to-end, taking a drug end-to-end.
Can you maybe talk us through where atom-wise sits in this very basic layman's model in
my mind?
Sure.
Absolutely.
There is a long thread of work and collaboration that happens to get the drug from a scientific
question to helping patients.
I think everybody gets into this work to help patients.
I think that's the core fundamental idea behind this.
The beginning might happen as a fundamental research problem.
Think about the proteins in your body as machines on an assembly line.
Every machine is taking in a very specific input, transforming it in a specific way,
handing it to the next machine down the line.
Out of the food you eat or the coffee you drink, your body is breaking down that food
and then building more you.
One of the ways the disease happens is when one of those machines breaks or when it goes
haywire.
Now, if you were wandering a factory floor and you saw a machine going haywire, you might
throw in a monkey wrench that the machine is busy chomping on the monkey wrench instead
of doing what it normally does.
Essentially, you shut down that runaway process.
In fact, that's the way that a lot of medicines work today is that they just physically block
up a protein from doing its normal job.
To go back to the assembly line metaphor, you'd want it to stick very well to that machine
to be able to shut it down.
You also want to make sure that it sticks specifically to that one machine and doesn't
shut down 100 other machines on the factory and cause all manner of side effects.
That's why being able to predict does it stick to what you want it to stick to and does it
bounce off and not impede what you want it not to impede to is so critical as a problem.
That's what we did.
We framed that as a machine learning problem.
We said run a prediction rather than put in a test tube because test tubes are expensive
and laborious and tricky instead of physical experimentation, run it as a prediction problem.
That's where the machine learning link comes in.
Are you using this predictive model to develop your own drugs or are you creating a platform
that allows other drug developers to use?
We're doing both.
We're doing both.
We have a portfolio and the portfolio has multiple different strategies in it.
We work with big pharma companies like Bayer and Eli Lilly.
We work with emerging biotechs like HANSO and Stamonix.
We work with a number of academic joint ventures where we've held academic spin-out companies.
They have biosciences, for example.
We have our own pipeline where we're building our own projects and doing some of that really
heavy lifting on our own.
I would love to get a better understanding of where your model first focuses on.
Are you taking a protein that you know is involved in some type of pathology and then
going through a long library of potential molecules that could help fit into that protein?
Or are you taking a drug or a molecule that you think could maybe be helpful and then
trying to see all the different pathologies that it could possibly help with?
So it's the first one.
So the way we approach it, we do something called structure-based drug design.
So we begin with the shape of the protein, and that's going to inform which molecules
we think will fit there.
So often it's our partners that come in with some protein or some idea and say, "So we
actually have one of the biggest academic collaborations in the world."
In the early days, we needed to prove that our technology worked.
People have been trying to use computers to do chemistry for decades.
It's a good idea.
If you think about it, Pharma is the last major industry where you build and test every
prototype physically.
When Airbus designs a new airplane, they will simulate a thousand wings before they ever
build one.
Wow.
That's really interesting.
Right?
And then you still build it and still, you still test it, right?
Like you still take it to the wind tunnel, you still do a test flight before I ever get
on it, I hope, right?
You do a thousand tests computationally, virtually, right, before you ever do a physical one.
And Pharma is really the last major industry where you build every single prototype physically,
bespoke manually, and then test it physically to see what it would do.
And so one of the ways about thinking from a business point of view, what we're doing
is we're bringing the efficiency that every other form manufacturing has had, we're trying
to bring that to the Pharma industry.
I'm curious to understand a little more about your technology.
So your technology, AdamNet, you have a paper on it available online linked to the website.
So the secret sauce is out for everyone in the world to view.
One, are you not worried about having that secret sauce out there?
And two, could you describe a little bit on what the technology is?
Yeah.
You want me to give the secret sauce on there?
I'll give you the two sentence version.
Probably many of your listeners are familiar with machine learning in the context of image
recognition.
So an image to a computer is a two-dimensional grid of pixels, and every pixel has red and
green and blue color channels.
Well, proteins are 3D.
So instead of a two-dimensional grid of pixels, we set up a three-dimensional grid of oxles.
And now every grid point, instead of having red or green or blue color channels, we set
up with oxygen, sulfur, nitrogen, SP2 carbon, SP3 carbon, et cetera, color channels.
As soon as we do that in coding, then you can take the algorithms that work in two dimensions
for image labeling, image understanding, image recognition, and we can translate those algorithms
over to the domain of biochemistry.
There's the secret sauce.
That's how it works.
So that's the foundation of our AdamNet technology.
Just to complete that analogy now, and when we're thinking of images on the output we
have often, is this a cat?
Is this a dog?
What is the equivalent way to think about that in AdamNet?
Yep.
So the simple way, you could set it up as a binary classification problem.
And so you could say, you could do a cutoff and you say, below one micromolar potency,
we're going to call that as an active.
Below 30 micromolar potency, we're going to call everything weaker than that as a negative.
And now we just have a binary classification, so you have a labeled training problem.
And I'm sorry, but what is the molar potency?
What is that metric?
This is a measure of the strength of the binding between a drug and a protein.
And then the label is a measure of if they bound or didn't bind, or a measure of how
strongly they bound or not.
I'd love to come back to this strategy of you have a very unique, very cutting edge technology,
and yet you've put it online for the world to be able to read about and possibly replicate.
So what gives you that confidence?
I think in the limit, if I can write a line of code, somebody else can write that line
of code.
So at the end of the day, it's unlikely to be a secret algorithm which provides your
long-term competitive advantage.
It's some business process which provides your competitive advantage.
In machine learning, one of the things which is absolutely critical is how often you get
to test your hypotheses.
Now often you get to test really prospectively test your models.
People talk a lot about the size of data sets of training data, but actually the number
of independent prospective tests, real test sets that you have, that I think is underappreciated
and absolutely critical.
I still review papers, submissions in this field and so forth, and all too often I get
papers which sound like this where somebody comes to you and imagine we're not doing chemistry
prediction, but imagine we're doing financial prediction in the domain where you want to
predict the stock market, and somebody comes to you and says, "Hey, I was able to predict
yesterday's stock price to within a dollar.
Give me your life savings.
We're going to be rich."
Probably you would say, "Okay, that's interesting.
Let me credit where credit is due.
That's interesting.
If you couldn't do that, certainly I wouldn't give you my life savings."
Have you ever predicted tomorrow's stock price?
That's a fundamentally different challenge, and it's much more convincing when you predict
tomorrow's stock price, and then tomorrow comes around and you see whether you're right
or wrong than predicting yesterday's.
True?
True.
In fact, what you need, the critical thing is, predict tomorrow's and predict tomorrow's.
You discover these kinds of patterns.
Either you're a genius to discover them from the very beginning, I am not a genius.
You have to discover these things prospectively by actually testing.
We wanted to demonstrate to ourselves, to ourselves, that this would work and it would
work robustly, that it works on a hundred different proteins, that it works on different
diseases, that it works in different people's hands, but I'm not an expert in a hundred
different protein.
We decided to partner with academics.
Imagine you're a professor, let's say at Stanford.
Imagine you believe that protein XYZ is the linchpin protein in Alzheimer's or in cancer
or in COVID.
You would come to us and you would say, "I want molecules for protein XYZ.
Go do your AI Hocus Pocus on protein XYZ."
We go and screen billions of compounds for academics, millions of compounds for commercial
projects, billions of compounds.
We then buy the best molecules.
We ship a 96-well plate and the academic can put them in their experiment and tell us where
they're succeeded.
To the academic, it's free.
It looks like a grant, if you have people who are listening, if you search for the artificial
intelligence molecular screen, the atom-wise Ames program, you can apply online and convince
yourself that this AI stuff works or doesn't work, why take my word for it?
Here we had over a thousand applications in the last 12 months.
I think from over 50 countries, we have been working on over 750 projects.
We have data back for more than 100 results.
Now you can look at this and you can really have confidence that our success rate, which
has been, I think the latest number was 75.7 percent of projects, we found something that
was exciting, our collaborator that they thought was a successful hit.
Now you can have confidence across this huge range of prospective results.
It's things like that, processes like that, which are a lot of work to put together that
give me confidence in the algorithms of confidence that we'll stay ahead of the game.
Every time we fail one of those, that's an opportunity to improve.
So by pouring more projects through the system, that's a chance because we build a single
global neural network.
We build a single global model.
Every project we work on is a chance to improve every other project.
That's how we stay ahead.
It's not about no one knowing the algorithm.
It sounds like what you're saying is that atom-wise is competitive advantage isn't necessarily
the technology and it's not necessarily access to the data or to some massive training dataset.
It's the partnerships that you're able to develop with people actually working on real
problems that you can start to learn about what actually happens in the real world when
you're running this model.
Does that sound right?
I think that sounds right.
I think, in fact, it's a D, all of the above, right, like kind of answer.
And so there is a technological advantage.
There is a data and a testing and a process advantage.
And then, of course, to build this stuff, you need a team which crosses multiple disciplines.
You can't do this just as a machine learning person.
You have to have the machine learning person saying next to the medicinal chemist who's
saying next to the software engineer who's sitting next to the structural biologist and
that they all care enough to learn how to talk to each other, right?
And they can't be siloed.
That's one of the challenges, I think, that big farmers have, that they have these people
but they're all siloed in different organizations.
You have to have them in one Zoom meeting all pulling together.
So with your technology, you are screening all these candidates and once in a while there's
going to be this bang, you've hit a target, there's this discovery at least from what
the computer is saying.
Could you talk a little bit about what happens then or I imagine that's already happened
a bunch of times and atom-wise and what the process from then looks like?
Sure.
Sure.
Let me give you a concrete example of one of our projects.
And so this is with one of our joint ventures called X37.
And if you go to X37.ai, you'll see a list of targets on which we're working.
And so the story I'd like to tell you is about one of those called PIM3.
So you have two proteins in your body, well, you have many more, but let's focus on two,
PIM1 and PIM3.
I'm not a biologist, right?
My background, I come from it, at it from the computer science side.
So my very high level view is if you block PIM3, then you will help cure people's cancer.
You turn off certain chemotherapy resistant pathways in the cancers.
These are stomach cancers, colorectal cancers, endodermal cancers.
But if you block PIM1, you will kill those same people by giving them heart attacks.
And so the whole challenge between this piece of biology, this corner biology, is how you
block PIM3 without blocking PIM1, how you have PIM3 selective compounds.
And as you can tell from the names, these are closely related proteins.
In fact, in the active site, that piece where you want the monkey wrench to fit and block
up, there's two amino acid difference.
So they're nearly identical.
So this was a target which lots of companies worked on, AstraZeneca, Roche, Janssen, I think,
Insight 6 or 7, big and small pharma companies worked on.
And they have compounds that they discovered.
They have compounds that are in the clinic, but all of them are non-selective.
All of them are called PAN PIM inhibitors.
So they're in the clinic, but they're struggling with unacceptable cardio talks, as I understand.
And so our collaborators said, look, if you could, if the AI could deliver selectivity,
that would be hugely valuable.
And so in this case, we screened 11 billion molecules.
And we pulled it down to fewer than 500 that we tested physically.
We had, oh, and I should mention that there is no structure for PIM3.
So designing is challenging here because the information is missing.
So actually we had to infer a structure.
We had to build what's called a homology model, which used PIM1 as a template and then put
in the edits.
So you can think of it as a simplified protein folding problem, maybe.
Protein folding with a lot of hints might be a way of thinking about it.
And if I can just ask a quick question here.
When you say there is no structure for PIM3, what do you mean by that?
Sure.
So people may have heard of the protein folding problem as one of these classic open problems,
one of the challenges of our time.
And so the idea is proteins are a long string of amino acids that then self-assemble.
They twist up and they fold up on themselves and they turn into these three-dimensional
nanoscale machines that run all of life's processes.
And so the question is, how do you go from the sequence, which is pretty easy to figure
out, to the 3D structure, which is pretty tricky.
And if you've ever done origami, right, or seen a master of origami, you know how intricate
you can get just by folding it on oneself.
I mean, it can be things of rare delight and beauty.
And that's with only some hundreds of years of humans thinking about it, right?
Like, now if you have billions of years and the most hardcore test, which is evolution,
the kind of intricacies that you could pull out.
So that shape, remember, right, what we're trying to do is we're trying to block up exactly
one of these proteins, not the others.
And so those tiny nuances, those tiny differences in shape are critically important to getting
a handle on how you would go about blocking PIM3 without blocking PIM1.
So today you can try to fold a protein computationally.
One of the things that the AI system actually allows us to do is it's much more robust.
It turns out that previous generations of computational approaches were pretty brittle.
If you got things not exactly right, they would go off the rails pretty fast.
If you got everything exactly set up, they could predict, but they weren't very forgiving
of errors.
And there's always errors in the data.
Everybody who's ever looked at data sets or generated data sets knows that there's error
bars on data.
And so empirically, it looks like machine learning is more robust and more forgiving
of errors in the input data.
And so we're able to use computationally inferred structures, computationally predicted structures,
what are called homology models.
That sounds fascinating.
We've been in suspense for the PIM3 discovery.
So I'd love for you to be able to continue that story and tell us where that went.
Okay, I'll give you the short version just because people are waiting in anticipation.
So where we last left our hairs.
So there's the PIM3 has utility in helping treat cancers.
PIM1 that causes heart attacks if you block it.
So the whole challenge was how you hit one without the other.
There was no structure of PIM3, and so designing was difficult.
There were no known selective compounds out there.
So kinds of machine learning that need molecules as a training set, not applicable here.
Computational tools that need an experimental structure, not applicable here.
Humans hadn't been able to crack this problem.
What I really like about this is this is a head-to-head comparison of the AI technique
versus all the other techniques, right, because Roche and AstraZeneca and Jensen, they have
all the other computational techniques, but they also have all the physical techniques,
all the physical experimentation, the high throughput screening, the DNA encode libraries,
the fragment-based, et cetera, et cetera, the human insight and so forth.
And so this was a case where we set up and we screened 11 billion molecules.
Which I have a comment on that once I'm done, by the way, we're living through a fundamental
transformation in the way chemistry is done right now, and that's, it's pretty exciting
to be able to see that.
But we screened 11 billion molecules.
We ended up physically testing fewer than 500.
36% of those had some activity in the assays, so it worked really very nicely, and seven
of those actually had triple-digit nanomores.
So better potency than you usually get out of physical experiment, which gives you a
whole range of different paths to start working down.
AI was able to deliver something that historically people had tried repeatedly, good people,
working hard, smart people had tried, and just weren't able with the previous tools.
And so I think that's really what we want.
We get excited when we can unlock new biology, and if you want to unlock biology, which has
never been possible before, you need tools that have never been available before, right?
Like that's, and that's the role of AI.
That sounds fascinating, and does sound a little bit like a superhero story there.
Now, I'm curious, you know, taking the superhero sort of analogy, you know, you have great
power, great responsibility, you come up with this molecule, do you, what do you do then?
Is this molecule now a secret sauce, which you can't share with the rest of the world
before you patent it?
And once that's done, what's next in terms of getting it out to the world?
Sure.
So, so it depends.
We work with, as I mentioned, right, we work with a lot of academics, and there's academics
who just want to publish, right, like they want, they want to do basic research.
That's what their motivation is.
And so that's fine, they can do that.
And then there are other researchers either through, through our big pharma collaborations,
or small pharma and emerging biotech joint ventures, or even with academics, or internally,
where we want to help cross that bridge and really impact patients' lives, right, like
really have new medicines for patients.
But there's this long stretch of really going through clinical trials and proving that these
things work.
And so those are the, those are the next steps, is getting it to that point where, where we
can move into the clinic and then through the clinics to the patient.
I'd love to talk about your efforts towards TB and malaria.
Now, when we think about drug discovery, there's a little bit of association of, these are
the big bad players, they're out to make money.
But TB and malaria have historically had not as much attention just because there is more
of a global health and a social, social justice component that's associated with that motive.
So I'm curious if you could talk a little bit about both your motive to get there and
the progress that you've made over the, over the past few years.
Absolutely.
If I, if I can phrase, rephrase the question, there's, you know, one of the things that
AI is, is discussing and thinking about vigorously right now at this moment in history is, is
fairness in AI and equity in AI.
And I think the version of that for health AI and pharmaceutical AI is, is it's a question,
I think what you're pointing out, which diseases get picked to work on, right, which research
gets followed up on.
And if you think about it, if you go to other places in the world, there's other disease
burdens, right?
There's, there's a different disease burden because there are different environmental
dietary genetic factors because of the, the history of, you know, history of the people.
And so if I remember correctly, in Southeast Asia, stomach cancers has a higher incidence
in East Asia, diabetes is more prevalent in South Asia, it has a large number of cardiovascular
and so like, there's nuance here that, that you have to care about.
In addition to infectious disease, the chagas and, and malaria, tuberculosis.
So I think, you know, one of the, one of the real advantages of new technologies like AI,
which allow discovery, which allow discovery to be done more efficiently with a higher
probability of success with, with shorter timelines is that you can go after diseases
which have more diseases than you could have before, right?
And through programs like our Apes program, we can access researchers around the world
and that's what we're doing.
I mean, we're empirically, that's, that's already what we've done.
I can give you an example and this is, this is a little bit just to, to push this idea
a little bit further.
There's something called rare diseases and ultra rare diseases, sometimes these are called
orphan diseases.
So one of the projects that we're working on, which we published on recently in, in journal
and medicinal chemistry is cannabine disease.
So cannabine disease is an ultra rare disorder.
It is genetic, if, if you're pregnant, it's one of the ones that you get screened for.
So this is, there's a molecule in your brain, it's called anacetyl aspartate, NAA.
You have a system in your brain that makes it and a system in your brain that breaks
it down, that degrades it.
And these kids lose the ability to clear it out, to degrade it.
And so this NAA builds up, the myelin sheath around the neuron starts to degrade.
The kids stop hitting developmental milestones and, and it can be fatal and there's really
no treatment.
So it's quite tragic.
Now, we had a hypothesis that these kids having lost the ability to clear NAA, if we then
slowed down production, we could bring it back into balance.
And so we had this hypothesis that we could go after the N acetyl aspartate synthetics,
the production side.
But this is one of these classic, undruggable proteins.
It's a membrane-associated protein, sticks to the, to the cell membrane.
And that made it very challenging to work with, very challenging to purify, very challenging
to express.
It was just hard to get enough of this protein to run what you might traditionally do, physical
experimentation.
And so people hadn't been able to do that approach.
And so that meant that there were no drug-like molecules, which were known.
So machine learning methods based on, on the drugs had no training data to get going.
Now our machine learning technology, one of the things that I mentioned is that it can
work when you infer the structure computationally, right?
So when you build a homology model.
And so that's what we did.
We actually picked as a, as a template, as a guide, a bacterial protein.
So understand, these diverged three billion years ago from the human.
I mean, like, these are very distantly related.
There was something like 17% sequence identity between the bacterial protein, the human protein.
And so this is the quite, quite distant.
And yet we decided to, to take a shot and go for it.
In this particular case, we, we tested 7.2 million molecules computationally.
So we screened 7.2 million.
That took us two hours.
These days we do 16 billion in, in about two days.
And we pulled it down and with our collaborator.
And so this was Professor Ron Biel at the University of Toledo physically tested 60 compounds.
Now understand that is low throughput, you know, grad student pipetting levels of throughput
that we could do.
And so how do you, how do you bridge that link?
Well, you bridge that link with, with AI, you have to make sure that those 60 have as
high probability as possible.
And so of the 65 of them were active.
The best one, all the drug, like the best ones actually 400 nanomolar potency, which
is better than you would get out of a physical screen.
So there is, is, you know, the distance between that and a drug is, is really still quite
far.
And so we're continuing to, there's more to do, but now we have a toehold for something
that had been intractable beforehand.
I think that's very exciting.
I know we're out of time here, so I'd just love to ask you one final question.
You mentioned earlier that AI is about to change the way we do chemistry.
I'd love to understand your view on five years from now.
What is the role for AI and chemistry and particularly drug discovery?
What is that world going to look like?
If you think back, like we're, we're living through, we're living through and have been
living through just a fundamental change in the way that, that chemistry can be done.
And why small pharma companies, there's, there's this beautiful flowering.
I think, I think the, the short answer to your question is, is I think there's going
to be an incredible flowering because we're going to put incredibly powerful tools in the
hands of researchers around the world.
And so we're going to see just an incredible opportunity whenever you've been able to,
to access the long tail of research, you find there's, there's amazing things out there.
And every time we've gotten to the long tail, we found amazing things out there.
And I don't care whether you're talking about Amazon with books or Netflix with documentaries,
right?
Like there's incredible things that you can discover out there, which, which weren't
available, just focusing on the most popular things.
But to access it, you fundamentally need new technologies and new business processes
enabled by those technologies.
And those always go together.
And so that's what we're going to see, putting these, these powerful, powerful tools into
the hands of people that are, that are studying all manner of diseases.
Fundamentally, I think that's what we need.
If you think about, you know, COVID, I mean, if I can take it back to the topic of the
day, right, like we have 15 projects on COVID, why the answer is nobody today knows what
the best, best treatment is going to be.
Nobody knows if it's vaccine or if it's antibody or if it's, it's drug repurposing or if it's
novel small molecule.
And the answer is we got to try all of them in parallel because science is messing in
the only way we figure out which one works is by trying.
And so if you want a shortened time, you got to try these things in parallel.
But that's true for every disease, right?
That's true for every approach to cancer and Alzheimer's into Parkinson's and fibrosis,
right?
And so you want to, you want to let a thousand flowers bloom.
So if you're interested in working on this, anybody out there, right?
Like we're hiring, but, but like that is, is the way that the farm industry will be working.
And what we will see, what we'll be able to measure is this incredible flowering of, of
discovery and approaches and who is doing discovery.
Abe, I have learned so much in this conversation.
Thank you for this, truly.
Thank you.
This, this was a lot of fun.
And that's all folks.
A big thank you to Abe Hi-Fits for talking to us today.
And thank you for listening.
We're your co-hosts, Pranav and Adriel, and until next time, stay safe and stay healthy.
The AI Health Podcast is produced and edited by Oishi Banerjee, music by Ethan A. Chee.
Many thanks to consulting producer Margaret Coucher.
If you like what you just heard, let a friend know.
Subscribe to this show and give us a five-star review on Apple podcasts.
Follow us on Spotify or connect with us on Twitter at AI Health Podcast.
Podcast Summary
Key Points:
Introduction to the AI Health podcast exploring AI in healthcare, biotech, and medicine.
Discussion on how AI is revolutionizing drug discovery in terms of cost and time.
Explanation of the drug discovery process, focusing on target identification and molecule design.
Overview of how machine learning, specifically hit discovery, is aiding in identifying potential drug candidates.
Atomwise's utilization of deep convolutional neural networks for drug discovery.
Importance of structure-based drug design and prediction in optimizing drug development.
Atomwise's collaboration with big pharma, biotechs, and academic institutions in drug development.
Description of AtomNet technology using three-dimensional grids and color channels for biochemistry applications.
Summary:
The AI Health podcast delves into the transformative impact of AI on healthcare, biotech, and medicine, with a focus on drug discovery. The episode discusses the escalating costs and timeframes involved in bringing new drugs to market, emphasizing the necessity for more efficient methods. It explains the drug discovery process, from target identification to molecule design, highlighting the role of machine learning in hit discovery.
Atomwise's pioneering use of deep convolutional neural networks in drug discovery is explored, along with its collaboration with various industry players. The AtomNet technology is unveiled, showcasing its innovative approach using three-dimensional grids and color channels for biochemistry applications. The podcast sheds light on the significance of predictive models in optimizing drug development and the need for prospective testing to validate hypotheses effectively.
FAQs
The average cost of getting a new drug to market has risen to about $2.6 billion.
Getting a new drug to market takes an estimated 10 to 15 years.
The goal is to find molecules that can bind to target proteins and modify them to treat diseases.
Machine learning helps identify molecules that can bind to a target protein and evaluate their activity.
Atomwise uses AI to computationally screen over 16 billion molecules to discover novel drugs.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.