Early-phase clinical trials differ significantly from late-phase ones in objectives, design, and statistical approaches. In early development, the primary goal is to efficiently decide whether to advance a compound—often with limited data—making it essential to use scientifically grounded assumptions, such as pharmacokinetic models or covariate adjustments, to extract maximum information and increase statistical power. In contrast, late-phase trials focus on regulatory approval and minimize risks like type I errors through non-parametric, robust methods. Dr. Brian Smith highlights that early-phase statisticians must balance scientific realism with statistical rigor, often relying on historical data or Bayesian methods, though with caution to ensure representativeness. Key modern challenges include the scarcity of data in rare disease trials, the replication crisis caused by post-hoc analysis and overreliance on p-values, and the integration of machine learning—though these methods are more suited to large datasets. Smith emphasizes that the core of effective biostatistics lies not just in mathematical methods but in building trust through relationships with scientists, understanding the science behind assumptions, and communicating uncertainty clearly. He advises aspiring biostatisticians to cultivate curiosity, embrace a philosophical foundation for their methods, and build credibility through genuine engagement rather than technical isolation. Ultimately, the most impactful work in drug development stems from combining strong statistical principles with deep scientific and human understanding.
What are the differences between early-faced trial design and lay-faced trial design?
What are some common assumptions for early trial designs, and what are some challenges
for statistics nowadays?
Dr. Brian Smith will answer all these questions for you.
Brian got his PhD in statistics from the University of Kentucky and has been actively involved
in promoting quantitative sciences in drug development.
He has a diverse background in biostatistics and has helped numerous positions in the pharmaceutical
industry in academia.
He currently holds a position of executive director of biostatistics and opartus.
Throughout his career, Brian has focused on early clinical development in the pharmaceutical
industry.
He has played an active role in the development and implementation of statistical methods
to improve efficiency and quality of drug development.
With his vast experience in expertise in biostatistics, Brian has become a respected leader in
field and continues to contribute to the advancement of quantitative sciences in drug development.
Let's dive into this episode and see what Brian shared with us.
Dr. Brian, welcome to biostatistics podcast.
Thank you for having me.
Can you start by telling us about your background and how we became interested in biostatistics?
I was always relatively good at math.
I think the first thing you should know is that I'm the first person for my family to go
to college.
It was good at math, I went to college, and everyone said, "You should be an engineer."
So I went to a small liberal arts school, center college in Kentucky, and I was a math
and physics major, and I got to about my junior year and realized that I hated physics.
I didn't want to be an engineer, but I didn't know what I was going to do.
I thought I was going to be a math teacher and a football coach, because I played football
in high school.
But I had a friend.
We didn't have one statistics course as a math major, and I'd take it, and I'd like
it.
I had a friend who was a year younger than me, and she said to me, "Lay now on was her
name."
She said, "You know, you can get a degree in statistics.
I had no idea."
None whatsoever.
It turned out that her dad was with David Allen.
He was the head of University of Kentucky's statistics program, and so that's how she
knew.
But I mean, it goes, and then I joined, and eventually got a PhD, and then started working
in what initially was pharmacokinetics mostly.
First, at a university, and then later, I moved in to join Eli Lilly.
This was about 1996, pretty good time, and I worked in their early development group.
From there, I spent nine years doing that, and from there, I went to Amgen in California,
and again, in early development, and so that takes me to, I mean, they were about nine
years again.
That's about 2014, at that point, I joined the Bardis, so, you know, I've been working this
entire time as a bio-satisfisha in early development in the pharmaceutical industry, so that's
my background, and that's how I kind of got into it, and I kind of fell into it to be
honest with you.
I know I think that I was going to be here when I was younger.
So that's a cool story, but there are so many other applications for statistics.
Why did you end up choosing the biosephysics track?
Why did I choose biosephysis?
I didn't do reality.
This was in 1994, and it was not easy to get a job as a statistician in general.
I remember that I sent out 100 CVs, and I had two job interviews, so, you know, I was
like many of us, I just wanted a job, so, you know, I ended up getting to University of
Louisville, and working in their kidney disease program, and when I learned there, my
boss, who was a pharmacist, and I was interested in preparing nonlinear meets-effect models to
neural networks for predicting drug concentration.
And, you know, what I learned about pharmacist genetics and everything led to everything
else, but the reason I got into it, I needed a job basically is what it came down to.
But once I got into it, right, and I started learning about the things that I was doing,
you know, I was hooked, if that makes sense.
Yeah, it makes sense.
So I guess the more the more you the job you work on, and the more interested you get.
Yeah, man.
Yeah.
Yeah.
You mentioned you work on early development in your early career.
I'm wondering, can you tell us a bit more about what exactly is early development?
Yeah, sure.
Sure, sure.
We work from first in Yemen, so when a compound first gets to people through proof of concept,
basically, phase 2A, and we also support everything that has to do with clinical pharmacology.
Clinical pharmacology is basically pharmacokinetics, and it's basically making decisions associated
with how to dose a medication.
So, you know, the early parts of drug development, what you're interested in is finding information
that makes you determine whether or not to go for it with the molecule or not.
In the later part, we do a host, most of them are healthy subjects, a host of studies,
who basically informed the drugs label, you know, I don't know.
I mean, you've seen, you've been told before with a medication problem where they'll take
this with food or take it with food.
That's determining clinical pharmacology studies.
Also, it might say, you might be told, they'll take this with medicine X, right?
That's because, you know, there's a drug interaction problem with this happening.
So, that sort of information we do as well, that's how we get that information.
And that way, it's mostly pharmacokinetics, and it's all in the drug, it ends up in the
drugs way, if that makes sense.
So.
Yeah.
I see.
Thank you.
I'm wondering.
Is that now?
Yeah, yeah.
Very much.
Very much.
Very much.
So, I'm just thinking maybe a lot of the audience, they're not familiar with the early
development of the drugs.
But that sounds very interesting.
What you'd do is think about just a couple of different things, because it's true that
probably in the pharmaceutical industry, I would say three-quarters of the statisticians
work in later development, Phase 2B, Phase 3 sorts of things, a small proportion works
in early development.
So I want you to consider a couple of things.
First of all, one out of ten molecules they hit first in Yemen will eventually be submitted.
That's industry average.
That means that 90% of molecules fail during drug development.
So if you think about it to a certain extent, that means, so when you're working in Phase
3, trying to get a compound approved, you're doing a study in order to convince a regular
regulatory agency to approve your drug.
Then largely your customer, then, at that point, is the regulatory agency.
In early development, we're trying to help the organization decide with this portfolio,
whether or not to advance molecules or not advance molecules.
So the customer is not the regulatory agency so much.
The customer is really the company itself, right?
So that changes things quite a bit.
The other thing is, and just to be honest, a Phase 3 study might cost $500 million to
do, or a Phase 1 study might cost $5 million to do.
If you're going to kill a compound, and most of them will be killed, when do you want
to do it?
And you obviously want to do it as early as possible.
So what that means is that our studies find a very nature that both ethically and for
budgetary reasons, right, have to be a lot smaller than a Phase 3 study.
So we have to be able to with limited sort of information.
and gather and be able to help our organization decide whether or not it makes sense to move
the compound forward or not.
Now, why is that important?
All this stuff adds up together.
You know, in Phase 3, I mean, you've already demonstrated, you think you've really got a molecule.
So, therefore, in our customers, the regulatory agency.
So, therefore, what they're really worried about, the regulatory agency is worried about,
is making a type 1 error or bias, right?
Those things are important in early development as well, however, right?
Trying to the power, the type 2 error elevates itself to a large extent.
So, statistics is different on its focus.
And what that means to a certain extent is that we need to try to take every piece of information we have
and maximize the amount of information that it contains.
And so, that, in general, leads to the statistical issues between late development and early development.
In late development, you're going to be more non-parametric in approach, less assumptions that you made, more robust, let's say.
In early development, you're going to be willing to make appropriate scientific assumptions,
which actually add power to what you're doing.
And I'll give you an idea of just a couple of them.
One is through modeling itself, right?
A model itself makes assumptions.
If that's a appropriate scientific model, it also has a lot more power.
Where if I'm making no assumptions, you know, if I'm making no assumptions, I can't even really model.
I can, in early development model, I can, in early development, use historical information for say the placebo group, right?
In a Bayesian sense, right?
I can use, that's an assumption that's being made.
I can, there are little things that you can do.
I can decide instead of using a normal distribution of taking a transformation and looking at a logon distribution.
I can decide if it's appropriate to do a crossover study.
Cross-over studies in general have a lot more information that you can get out of them, just from that design choice alone.
Right now, I'm working on a project right now to demonstrate to our clinicians why it's such a bad idea to dichotomize a continuous variable.
That is, take a nice continuous variable and choose a point of whether or not someone's a responder or not responder.
It turns out to a certain extent, and it depends upon what assumptions you make about what's generating the data.
But the power for seeing whether or not the proportions are different between two groups, right, is a lot lower than it's seen the means between the two distributions are different.
In fact, you would need something between 50% to 100% more subjects most of the time to be looking at the proportion.
Now, we have limited data to start out with, right? That's a way. So in general, any time an analysis choice can give you a lot more power with relatively reasonable assumptions, we should do that to get the most information of the data.
And that's really what, as a statistician, I think, concerned with to a great deal, written some work on it and so forth and so on.
So does that make sense? The big, the big difference there in mindset.
I'm wondering when people talk about making assumptions, what is the boundary? Like what is the limit of how, I guess, how much of a assumption you should make?
Because I feel like we know, for instance, that in this just has how drugs work. You give a dose of drug that it's not the dose that causes the effect, it's actually the concentration of the drug.
So you and I are different, right? Because we're ethically different, because we're different genders, right? We, if we take the same drug, we're likely to have different concentrations.
So, which means, if, whatever reason, the drug that we're taking is the one that I, that I have higher concentration, what can you do? And we take the same dose, what that means is two things.
One, it's more likely that I'm going to have a positive effect. And it's also more likely that I'm going to have a toxicity, right?
So, so in general, and we know other things about how drugs work, right? They hit a, they tend to hit a target. So you take a drug, it hits a target.
There are a bunch of targets floating around in your system. There's only so many. There's only a maximum number of targets that you can fit to, to affect the system.
That in essence basically causes a plateau to have the widows or concentration of that. That's a reasonable assumption, right? That's, that's basically biologically, that's how drugs work.
So, therefore, using that assumption is a reasonable one. If that makes any sense. So, that's, that's what I'm talking about. Use the science. Use your scientists to give you a reasonable idea of things that you can do.
Yeah, that makes a lot of sense. Thank you. So, all of us, most of the assumptions you're talking about are clinical assumptions?
No, I mean, they turn into a model. So, so it's basically clinical assumptions that turn into some sort of mathematical model, right? And that, that's part of it. I mean, I think the other part of it is just simple things.
For whatever reason, we are used to doing linear regression. So, we always do it. It's oftentimes actually very bad assumption. Right? We're oftentimes, for whatever reason, we're stuck on the normal distribution. And so, or we're stuck on the mean.
Until we don't transform what we should, right? And we, we know, for instance, most of the data that we collect is continuous and positive, let's say, right?
A normal distribution assumes there's positive probability for values that are negative. And yet our data itself can't be negative.
We use a distribution that you know a priori is not the right distribution. By law transforming it, you use a distribution that basically fits the sample space of your data.
And it turns out you're going to have more power when you do that. Right?
We always can rely if we have a large sample size on the central limit.
But we don't have the luxury in early development to rely on the central limit.
If that makes sense. So, therefore, you know, a smart transformation, you know, makes a lot of, makes a lot of sense.
Right? And, you know, that's simple things. Like, I mean, using covariates, right, in your model that always improves, especially using the, if you have a baseline value that you're using as a covariate, that always improves your precision of your estimate.
So, if that's the case, then use it. Right? Don't ignore that. That's sort of information. So things like that. Right?
I see. That makes a lot of sense. I'm also wondering when you're talking about late development, you're saying it's sort of a little bit more nonparametric. You make less assumptions.
I'm wondering, can you elaborate a little bit more on time? Well, I mean, because, I mean, you've got to, you've got to think from it, and I'm not a regular.
I've done some of the years, but you've got to think when you're pointing to it. Right?
I mean, the thing that they don't want to happen is compound to go forward because of you assumptions that you made that were turned out to be wrong.
And that that compound gets the market and, you know, gets sold to people, and it really doesn't work.
Their tolerance for a type 1 error or for an error happening due to assumptions is very hot. They don't want that to happen, right?
And say you are, in essence, when you're thinking about statistic way, how to put something together, you have to take that consideration in the case.
In early development, again, the customer is not the regulator. We're trying to make a good decision right in any particular point in time, right?
So therefore, you can make more assumptions. In fact, I think, really, if you think about what I've kind of discovered, you always hear basing in frequencies targeting with each other.
And I don't think that's the issue. I think really the issue is for the problem that you have, how many assumptions can you make? Right? Is it reasonable to make?
Before you mentioned probably sometimes people use Bayesian method to incorporate the prior information to the placebo group. And do you think they use it more in the earlier development phase?
And that's also because of the assumption making statement? Absolutely, because you are, I mean, you know, that's one of the things you got to be really, really careful about, because when you're using historical information, I mean, every study that you have has an inclusion exclusion criteria and has patients that you're bringing into the study, if the historical placebo information is not representative of this new class of individuals, then it could actually be very harmful.
You know, it could cause you to make mistakes, right? So you have to be really, really, really careful to try to get likes and likes matched up together, right?
So, you know, to me, that's one of the largest assumptions, you know, we do it every day, we have to be really careful when and if we do that. Obviously, you know, unless you're in a situation, there are exceptions that we rule, I mean, unless you're in a situation for really grievous, you know, bad disease and a really rare population, right?
As a regulator, you would be, I would think very suspect of that, you know, but everything, everything in life is dependent upon what the disease is, you know, how horrible a disease that is, how rare the disease it is and so forth and so on.
So, I see. Thank you for sharing that inside. It's really interesting because I don't think I've ever considered early phase trial and later, later phase trial in this way, because for sure, obviously, I haven't experienced a lot of, I haven't had a lot of experience on them, but that's very interesting. Thank you.
Welcome. I'm monitoring throughout your career, because you've been in the field for a long time. What is the most, I guess, interesting project that you've worked on or like the most interesting submission that you've worked on?
I think one of the, one of the more interesting things that I worked on and this gets back a while, was around, I don't know, the early 2000s, there was a guide document, ICHE 14, for the analysis of PTNable, and it turns out that the QTNable was something that, say in the 80s and 90s,
wasn't necessarily studied that much by drug developers, but there turned out to be one or two instances of a compound that gets something called the Earth channel.
It ended up causing people to have something called Twisaud de Ponce, which is an arrhythmia in key and cause instant death.
And this is for molecule, the one molecule for this happened with was a anti-histamine. I mean, it was being taken by individuals that basically had allergies.
You don't expect when you're taking a medicine for allergies that you could potentially die, right?
So this happened, like I think this happened in the early 90s, yeah, I think the early 90s.
And regulators were concerned and they started looking at it and they came up with, you know, eventually guidance documents.
At the time, I was working on a compound that had an impact on the cardiovascular system, which made it potentially look like there was a QT problem, but it turned out there's a correlation between the QTNable and heart rate.
The compound was causing the heart rate to increase. And what looked like a QT problem was really due to heart rate increasing. And so we came up with my co-author and I, Alex Michealanka, came up with a method for analyzing QTNable, that accounts for the heart rate.
I think in a relatively interesting fashion, it's basically uses heart rate itself as a covariate controls for it, so that QT becomes basically because of the analysis and intended to partrate itself.
So that was very interesting. I mean, the other things that I've always been interested are associated with, you know, how to get the bang.
You know, how to get the bang for the buck out of the analysis that you're doing. Right, how to get the most information out of the analysis that you're doing. So I've looked at different, different components associated with that and published on it.
So, I don't know, those are, those are just a couple, couple different, different examples.
Okay, share more life on the maximizing, I guess, optimizing the information you can use for analysis.
And then it comes down to that assumption, right? You know, I think, and I don't think we communicate it very well.
I mean, one of the, one of the things we sometimes have an analysis, analysis A that's better than analysis B.
Well, what, one of the things that happens in Madison in general is, I almost feel as if there's the first person that does the analysis dictates what the analysis was always going to be there on after.
So in 1983, I have a compound at doing analysis, and everybody else tries to do the same thing to match the results.
Well, if that's an analysis is an optimal, basically.
Then, in essence, you're doing a non-optible analysis over and over and over again.
So, in general, we, especially in early development, we have the opportunity to do analyses that are not standard, that aren't the ones that everyone's always done in the past with that end point.
I see. Thank you. So I guess another thing I want, I'm wondering is since you've been in the field for a long time, what are some of the development that you've seen throughout the past few years.
And then what do you think are some other new developments or innovation you can see in the future?
Yeah, I mean, I think the number one development has happened through the years. I'm not sure the basic tenets of statistics has changed very much, but what's changed is computing power.
Computing power enables you to do things that say when I started in the 90s, or when people worked in the 50s, couldn't even dream of doing.
So, and yet, if you look at the things that we do, you could almost say that it was envisioned by things that Fisher did or by what Savage did back really early on.
It's just they couldn't do it, they couldn't computationally do the things that their methods that they had could do, if that makes any sense.
I think what the future brings us and thought about this a little bit, I mean, I think there are three components to think about for statisticians in the future, or three problems we have to think about, big data, small data, and in the replication crisis.
Let me talk about about each one of those separately. We'll know with big data.
This has been an area that data science primarily is open to to attacking two machines.
learning methods. It turns out from my perspective that machine learning methods are very non-parametric
nature. They tend not to make very many assumptions. And that's the reason with large data sets they
work. It's hard to apply machine learning algorithms to small data sets in general. So as we get
are getting in places with more data, obviously we can use that. And one of the problems is
is that in drug development we're doing clinical trials. So we are in essence handicap to a certain
extent about how much data we can even get even in our largest phase three studies. It's nothing like
what Google gets. But one of the places that I do think that we have to think about is that with
wearable sort of devices, we can get lots of information for a subject. That's different than
having lots of information from many subjects. But I think that's one area that's not really thought
what do you do if you have if you're measuring something every 30 seconds or six months?
In a subject, obviously, maybe it's not obvious, but obviously there's still variability subject
to subject. So I can't just get an infinite number of observations from a particular subject
and know everything, because there's still variability that's left over from subject to subject,
but certainly having that lots of information on to make us understand what's going on with that
person better and remove part of the variability. But how to deal with that? I don't think it's
good. Well thought out at this point to a large extent. All right, small data. The reason that we're
not just going to have small data in early development, but one of the interesting things that's happening
in drug development is we're seeing more and more drugs for rare diseases.
Now the problem, that's good. And we're seeing in oncology, for instance, a breakdown of
studying drugs and smaller and smaller portions of individuals that have a disease due to genetic
factor. What that means, though, is the number of subjects that you can potentially recruit because
there's more and more limit. And if I have less people with a limit is less for data than I do.
So I think we have to think about maybe some of the principles that I talked about in early
development have to apply to the situations of rare diseases. Right, and we have to think about
that more. And I think we'll see more of that. And the last thing is, you know, back, I'm not
quite sure when this happened. Maybe some people say it was fissure, but it was decided that,
you know, the p-value is less than 0.05. It meant you had something that was greater than 0.05.
You didn't. I don't think statisticians felt that way, but I think the scientific community did.
And I think that we've seen relatively recently, you know, how dangerous that could be.
So in general, we have to do a better job of talking to our scientific colleagues
about what this information actually means. And things like multiplicity and other sorts of
things have to take takes center stage. So because otherwise, with the sort of sort of failures that
we've seen, it makes our profession look bad, even though it really doesn't have anything to do
the profession has to do with misunderstanding, decision making in general. Right, so these are,
I think, our biggest concerns that we need to address.
How was the, I guess, the type when our debate related to the replication crisis that you were talking about?
Well, I mean, so in general speaking, right, I think the replication crisis is associated with
individuals, you know, many people that have worked with civic clinical people, right?
They've seen a p-value less than 0.05 and that's on a subgroup analysis. It's time to go publish.
But the thing is, is that, you know, that p-value only is really a probability when it's answering
one question and I set it up beforehand and I have the thing power correctly and so forth and so on.
Just the p-value less than 0.05, you know, doesn't necessarily mean that that's not a type one error.
I actually describe it to a certain extent, a class by teacher upon this, sort of this way.
There are all kinds of scientific hypothesis out there. Think about scientific hypothesis where you
can look for them, right? There are coal mines and there are gold mines. If I'm going to mine for gold,
I'd rather mine than the gold mine than the coal mine. I'm afraid that a lot of times, right?
The p-value comes from analysis that's done after the fact and there's no scientific reason why
we did the analysis to start out with, but we have a small p-value. Under that circumstance,
you're mining for gold and a coal mine, right? And what I mean by that is if I get a set of
high hypothesis and all of them are false, right? We know that 5% of them are of p-value less than 0.05.
So what we see in the literature is only the things that are less than 0.05.
So everything is dependent upon where you start. Are you looking at something as reasonable to look at?
Right, so does that help? Yeah, but I'm wondering how do you think we as a scientific
community can avoid this kind of problem from happening? Well, I think the number one thing,
statistical education, which I don't think that we do a very good job of,
it's a statistical education with the people that we work with, so they understand
probability. They understand the caveats associated with probability and that a key value
itself is a conditional probability, right? And you know, this is where Bayesian sort of thinking
helps, right? Whether or not something is true or not is dependent upon not the p-value,
the p-value is just telling you basically, if there's nothing going on, how likely is it to see
the results that I got? If they're not very likely than the p-value as well, that's great. But the
thing is, is that before you even start there, there's some sort of, depending upon where you're
mining from gold, right? There's some sort of probability that there's something going on or not,
and so using that helps us understand things. And I don't think we need to do that formally
necessarily, but we need to have the concept that, you know, I need to start in a place
in which the hypothesis isn't bogus, right? And then at that point, I have more trust in the result.
So there's all kinds of information that I'd like to know when I see a p-value and that information
really is about how likely is this scientific hypothesis before you started and where did this come
from? And how many things did you look at? All of those things inform me how much, how much trust
I can put in the thing that I have, right? And I think we need to be better at communicating that
with the individuals that we work with so that they communicate in the manuscripts that they write.
That was a lot of sense. Thank you. Because I do think sometimes when I do physical analysis with,
I guess, none statistical people. All they want to see is the p-value less than 0.05. And when they
don't see that happening, they might switch the questions, they might switch the variables
including the model, which is, I don't think that kind of post-talk thinking is correct in a way.
Oh, it's clearly not. And that, you know, but
But I mean, I think that the thing is, is that you understand it, right?
I mean, people are, you know, we have pressure, I knew you were at Toronto, but people have
pressure, they need to publish things.
And that's where that comes from, basically, but, you know, it's not going to lead to anything
positive.
Now, on the other hand, you know, if I do something post-doc, and I see a p value of 0.00,
0.001, more than likely I'm going to have to plead to that.
Right.
Right?
So it really comes down to, you know, the strength of what the evidence is, the problem
is that we've got to this place where 0.05 is the gatekeeper, yes, versus that.
And it's not, it shouldn't be, right?
It's all within context.
So.
Thank you.
Very important conversation, need to be addressed better.
So I guess next, I'm going to ask a question on behalf of me and my peers.
So for people who wish to develop a career in Biostatistics, what advice do you want to
give them?
You know, I think that, I think the number one thing is to be really curious about everything.
Right?
You know, there are, there's lots of cool math, and there's also lots of cool science.
And it's understanding all of that together is really important.
But I think there's one other thing to keep in mind, and I don't think it's something
that's thought about enough.
So let me explain.
If you think about a statistician's job, there are two times.
When we work with a client where they really don't want to hear what we have to say.
And the one time is when we tell them, you can't do that study.
You need three times as many subjects.
That's not a, that's not a good thing, right?
That's a budgetary problem.
The other thing that we tell them is that, that thing that you really thought was going
to work well, it didn't.
Right, we deliver, we have to deliver those sorts of messages, both of them.
You know, I want you to consider something else and, you know, I've been, I've literally
been at parties before where I've introduced myself, they say, what did you do for a living
and they say I'm a statistician?
And they, you know, they politely walk someplace else because they think we're boring, right?
And sometimes are, but, but, you know, we, and hopefully this is getting better, but
at the university level, people's experience with statistics is a basic statistics class
where they're given lots of formulas, they're not really given much thought into anything.
And sometimes, you know, the teachers are even very good.
And so, in general, those individuals that are taking that class very much become the
scientists that I work with in the future, that is we sometimes start with a deficit.
There's not a high opinion about what we have to give and offer, right?
We have to somehow figure out how to overcome that deficit.
And it's clear to me that one of the ways that you can do that is by building relationships
with the people that you work with.
By building relationships with the people that you work with, it's kind of like this sort
of thing, right?
A friend, if they have bad news to give you, you still don't like it, but you might accept
that, right?
So, in essence, the de-agreate statistic, you need to be really good at that, yeah.
Do you need to understand the science, yeah?
But the thing is, is all that's meaningless, yeah, you don't connect with the people that
you're working with, right?
So, ultimately speaking, we have to work really, really hard in those relationships.
And those things don't, that work doesn't come at the computer.
That comes from coffee with a colleague that comes from, right, getting to know someone,
knowing someone's kids, knowing what they did over the weekend.
All those things basically lead to you being able to have a connection, which then allows
you some ability to have some credibility with the person that you're working with.
Because it's credibility that's really important.
Because if you have no credibility, and you've probably seen this, right, the scientists
that you are working with can just ignore you, right?
You've got to build the credibility, or otherwise you're not going to have it in there.
So, that seems to me really, really important.
And I think it's something that, I mean, I understand why in graduate schools, that
isn't necessarily taught.
But it's something that most people coming out of graduate school have no idea.
They have to work with someone, and they really what they have to do is create relationships
with individuals.
And each individual is different.
And each individual has different things that you need to connect with.
Right?
There's just no one size fits all.
So, I sometimes think we become, you know, being able to read people and understand
people is almost as important as the math that we do.
For sure.
Thank you.
That's a very good advice to give, because I think definitely communication and connection
with people is very important part of the work.
I guess that brings us to the last question, or it doesn't have to be a question.
But my question is, what is one question that you wish I would have asked, and how would
you have answered it?
And if you can think of a question, you can just share whatever you think that you want
to share.
But you have to.
Sure.
Absolutely.
And one question you did ask is, what's your statistical philosophy?
And the reason I say this, I think it's really important for everyone, not to have a statistical
philosophy when they first start.
I always love it when I've met a new graduate student, you know, they had a professor who
was a Bayesian, and so they did a Bayesian project, and they say, I'm a Bayesian, right?
But they don't really know what that means exactly, right?
I think that it's important to, statistics to a certain extent is a very philosophical area.
And there are different philosophical starting places that understand all of those and choose
one.
And what that means is, is that I believe that the information of the data is contained
in the likelihood function, and that most methods that take advantage of the likelihood function
are in the right area.
I also believe to believe that the distribution that you're working with matters, so we should
use distributions that are consistent with the data that we have.
These are two philosophies that I acquired, I didn't have them one day one, but I acquired
over time, if that makes any sense.
And I think that that's a really important thing for every young statistician to ask
them, so what am I, am I just a person that does methods, that can't be it, right?
There has to be a reason, besides somebody else recommending you to do it, there has to
be a reason why you do the methods that you do, right?
So have some sort of philosophical basis for that.
That's very interesting way to put it, I'm wondering when you're saying your likelihood
behave, sorry, believer, or you prefer consistent distributions, you're not exactly categorizing
yourself into the fricking disurbation.
No, no, no, no, no, because the likelihood function plays a very big role in Bayesian
statistics, and in fact, I think many, you know, just the side point, I think many of
the arguments between Bayesian and Frequence's are silly, because the fact is, under most
circumstances, if I start with a non-informative prior, right, if I start with a non-informative
prior inference that I get from Frequence's methods, and Bayesian methods, it's almost
identical, in the oftentimes incidentally, right?
The only thing is different is the interpretation, right?
Okay, let's put that to a side for a second.
the advantage of raising your hands.
statistics to a certain extent is the ability to easily use prior information right but you can
do that actually with frequency statistics as well you're violating frequency statistics but
but you can do that as well right if you're not using prior information though then the other
advantage of Bayesian statistics is there's cool software that you can use that can solve many
problems but if you're using a non-informed prior I would claim that all you're getting is a maximum
likelihood estimate that's fair right so so in general speaking I don't see where the big
deal is if when I'm talking to a client I read them about whether or not they're more comfortable
with the frequency of interpretation of the results or Bayesian interpretation and then I give it
to them that way because I don't I don't think it really it doesn't it doesn't impact action
right yeah it doesn't impact action so like I I just don't I so I'm not I'm not either one but
I do think the likelihood functions really important that makes sense it's just that's a very
interesting way to put it because I think after way too many arguments between fricantists and
facing them I know I like you know I went I honestly speak I don't think it was but I was
when I took my first influence course my my inference teacher was a militant frequentist and
this is a long time ago and we had a section in my inference course on Bayesian statistics and
the professor spent a whole lecture the day before they introduced this telling us why you know
Bayesian statistics was the worst thing in the world and you should never use it right right and
you know at the time I was like okay yeah that that's that's right but as a grown older I realized
I guess unless you have some deeply deep opinions regarding the underlying philosophy regarding
it but then it's a philosophical thing and I had to have a philosophical argument but the
reality is is I've watched statisticians right if it's not impacting what the action that comes
from the data it's collected you know you have a beer have the philosophical discussion and so
forth you can even yell at each other I don't care but the thing is is that the results the same
that makes sense yeah thank you so much for sharing all these great insights with us
and it's great talking to you thank you it was great talking to you too thank you for having me
and um I hope you have a wonderful day thank you youtube
thanks for tuning in and we'll see you in the next episode
you
you
Podcast Summary
Key Points:
Early-phase trial designs focus on making efficient, data-driven decisions to advance or discard drug candidates, unlike late-phase trials which aim to convince regulators with robust, definitive evidence.
Early development requires more scientific assumptions (e.g., pharmacokinetic models, covariate use, transformations) to maximize power from limited data, while late-phase trials prioritize non-parametric, robust methods to avoid false positives.
Modern challenges in biostatistics include managing small sample sizes in rare disease studies, navigating the replication crisis due to post-hoc analysis, and integrating machine learning with traditional methods in a way that maintains scientific validity.
Summary:
Early-phase clinical trials differ significantly from late-phase ones in objectives, design, and statistical approaches. In early development, the primary goal is to efficiently decide whether to advance a compound—often with limited data—making it essential to use scientifically grounded assumptions, such as pharmacokinetic models or covariate adjustments, to extract maximum information and increase statistical power. In contrast, late-phase trials focus on regulatory approval and minimize risks like type I errors through non-parametric, robust methods.
Dr. Brian Smith highlights that early-phase statisticians must balance scientific realism with statistical rigor, often relying on historical data or Bayesian methods, though with caution to ensure representativeness. Key modern challenges include the scarcity of data in rare disease trials, the replication crisis caused by post-hoc analysis and overreliance on p-values, and the integration of machine learning—though these methods are more suited to large datasets.
Smith emphasizes that the core of effective biostatistics lies not just in mathematical methods but in building trust through relationships with scientists, understanding the science behind assumptions, and communicating uncertainty clearly. He advises aspiring biostatisticians to cultivate curiosity, embrace a philosophical foundation for their methods, and build credibility through genuine engagement rather than technical isolation. Ultimately, the most impactful work in drug development stems from combining strong statistical principles with deep scientific and human understanding.
FAQs
Early-phase trials focus on determining whether a drug is safe and shows promising activity, using smaller, more cost-effective studies. Late-phase trials aim to confirm efficacy and safety for regulatory approval, involving larger, more complex studies with stricter statistical requirements.
In early development, statisticians make reasonable, scientifically grounded assumptions to maximize power with limited data. In late development, more conservative, non-parametric approaches are used to avoid Type I errors, as the stakes are higher and the data is more mature.
Early-phase trials emphasize using transformations (like log-scale), covariates, and historical data (e.g., Bayesian priors) to extract maximum information from small datasets, while avoiding dichotomization of continuous variables to maintain statistical power.
In early development, the customer is the pharmaceutical company deciding whether to advance a molecule, not a regulatory agency. This shifts the focus from regulatory compliance to portfolio decision-making and cost efficiency.
Bayesian methods allow incorporation of prior knowledge (e.g., historical placebo data) to improve efficiency and power in small studies, especially when data is limited and assumptions about the data distribution are reasonable and justified.
Key challenges include managing small data in rare disease trials, interpreting p-values in the context of multiple testing, and addressing the replication crisis through better statistical communication and scientific rigor.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.