Go back

The Evolving Role of Generative AI in Pharma

33m 8s

The Evolving Role of Generative AI in Pharma

In this podcast episode, host Alexander Schacht interviews Manuel, an AI specialist with a background in molecular biology and medical affairs at Sanofi. Manuel explains his career shift to AI, driven by a need to engage in technical discussions and innovate patient care. He describes his intensive training in Barcelona, emphasizing hands-on coding to deeply understand algorithms. Manuel highlights that generative AI is best suited for repetitive, structured tasks in pharmaceuticals, such as generating clinical reports and medical writing, where data formats are consistent. However, challenges include model hallucinations and strict data governance, particularly under EU regulations that restrict data transfer. The discussion also covers AI's role in coding, where it accelerates development, debugging, and research by providing instant code generation and analysis. Additionally, AI aids in data quality control, such as comparing outputs from SAS and R, and improves data visualization efficiency. Overall, AI is seen as a tool to automate mundane tasks, allowing professionals to focus on higher-value work that requires human intelligence.

Transcription

4772 Words, 25348 Characters

English
You are listening to the Effective Statition Podcast, the weekly podcast with Alexander Schacht and Benjamin Piscat designed to help you reach a potential lead grade science and serve patients while having a great work life balance. In addition to our premium courses on the Effective Statition Act Academy, we also have lots of free resources for you across all kinds of different topics within that Academy. Head over to zeeffectivestatition.com and find the Academy and much more for you to become an effective statistician. I'm producing this podcast in association with PSI, a community dedicated to leading and promoting you for statistics within P.H's Candles Tree by Benefit Authentication. Join PSI today for those that develop your statistical capability with access to the ever-growing and your demand content library, pre-reface tracing for all PSI webinars. Head over to zeeffectivestatition.com and find the more about PSI activities in the chamberpile side number 2. Welcome to another episode of the Effective Statition. I'm super happy to have a colleague here today because he is working all around Artificial Intelligence AI and many of us have been exposed to this topic already and it's definitely not a new topic, but it's still a very hot topic and a very fast evolving topic. I have certainly have used it here and there, but it's definitely something different to talk with someone that is working on it all the time. Manuel, welcome to show. Thank you so much Alex for having me. This is such a pleasure to be here. I hope I know the black sheep of your podcast because everyone is a statistician. I know that this is from an engineering point of view, so I hope that's near to a table. I definitely had non statisticians as well on the show, so you're not the first one. It should use yourself so that people understand where you're coming from and what you're currently working on. Perfect. I would like to say that I as a professional, we can say I am a hybrid because I am a molecular biologist, geneticist, what is my core training and I work on diagnostics in a hospital at the beginning and then in a pharmaceutical company, Sanofi, in the space of medical affairs for the disease. So we can say that my career started more or less in the medical science space, but then I migrated to AI and AI now we can say the biggest part of my focus because the thing was that when I was working at Sanofi, we started having some very small projects on AI, specifically for some therapeutic areas where we needed to bring innovation to make the patient journey a little more efficient and more different from the one that we had in the past. So in that moment, I realized first that I didn't know practically anything on AI and that made it really hard for me to have conversations with vendors on technical discussions because AI really be grasped of the field, but when we started discussing algorithms and accuracy and images and all of the things that for example, you need to build AI solutions, I was really lost and I was deeply passionate about the field. So I said, I think my past needs to take this other direction. So that's when I moved to Barcelona to study, Barcelona has one of the nicest schools of engineering in Europe and also Barcelona has one of the European super computers. So you have both of these really big spaces in AI in the same place. So that is why I moved there and I started my training in AI. It was one of the toughest things that I have done because I didn't come from the computer science space. So for me, learning was very hard. I remember that I spent almost 15 hours a day studying and practicing coding skills because I actually knew very little on how to code and the program required a very high coding skills development. So I needed to put my hands on the dirt and do everything there. It was very tough, but it was extremely rewarding because, thank God, it was the face released to the boom of chativity so that meant that for example, for you to have an idea, we need to do all the coding ourselves in and every student needed to write their own code. I don't know if I like that approach, but we can say that it worked, that they wanted us to think about the mathematical processes behind functions. So they asked us, it was actually a requirement that you couldn't use functions built by package. That means that for example, we needed to program the function ourselves so that they saw inside of that code that we understood the mathematical process behind. So we program stupid vector machines, very rudimentary and small to for you to encourage you to think about what you were writing. And they were really like focused for us not to copy the work of other students. So the code needed to be like developed by you and not a copy of other code. So it was actually very tough, but it gave me this mental exercise of thinking the solution before going and coding it, like how the data is going to be in the input space, how it's going to be transformed, how the output should look like, for example, what could be potential back in the middle of that function. If you have, for example, three dimensions here and then the output has four dimensions. So it was, I think, very challenging to learn, but it was very efficient. I think now I don't know how they are going to do with generation because you can generate a task on an exercise in seconds. Just to finish and wrap it up, I worked in several consulting companies. I also consulted for other pharma, especially on this journey of starting using AI for medical affairs first and then co-generation in for other companies. And now I am at Citel as a partner, I would say, into bringing all of these solutions that we have internally and to build a future where AI can be analyzed for people and to help us in some sort of way to take all the automatic repetitive tasks that we do day by day that would put us into the burnout space and actually allow us humans to, once we have the data and once we have perhaps the things that we need to analyze to use our human intelligence in ways that the language models cannot at this point or in developing it. I completely agree. It's yet another way of how to get rid of some of the in tasks and focus on those tasks where we have much more opportunity to add value. Let's speak about, where do you see AI being really already spend that used and broadly being accepted within the pharma space? That is a very nice question. In the sense that to start like defining this kind of space, the best use cases that we have for generative AI and agentic AI in these days is where you have something that you repeat day after day that involves the same data that enters the same data that goes out, the same changes in format, that is the best place for AI to come and help in that process. For example, we have a lot of documents that we need to build. For example, in the clinical development phase, for example, clinical reports that you need to produce, that are basically understanding data sets and understanding perhaps another report or another protocol, for example, that has the data but in a different format that then you need to grab that data and transform that into a new report that are the cases for generative AI because it's basically the data is there. The model just needs to understand the structure of the data, how it is being used and especially the context in which the data appears so that then the data can be transformed into a different format, understanding the context, the original context so that we don't perhaps use different words and change the meaning in your new document that we are generating. So when you find a repetitive task that you are doing over and over again, that is the case. Where do you see cutting edge areas where AI is maybe used here and there but it's definitely not yet widely adopted. I would say in the space of medical writing, we have a lot of opportunity in the sense that we are slowly approaching that phase and we have a lot of opportunity there because medical writing is one of the things that is being used. It takes a lot of people and it takes a lot of processes that they are mostly the same. The documents, the structure is fixed, that means that you are not going to have every part that you need for example for regulatory purposes, being changed like day after day. It's the same report, it's the same data that needs to be filled inside. The input is also the same. The problem that I'm seeing now that of course we need to address that also is the hallucinations problem that models because of the nature of how they are trained and how they see the world, we can say, it's deterministic. That means that you have a distribution of things that the model sees and that is the representation of the world for that model. Where you go out of that distribution, then we are subjected to basically randomness. That means that the output could be a good one or the output could just be outside of the definition of true that we have for that answer because in repetitive tasks sometimes you can more or less narrow the deterministic part because there are things that are being repeated step after step but it could happen sometimes that the model sees something that doesn't have all the knowledge to understand and then the output is not something that is valid. So hallucinations is something that we need to first of all define then understand and then learn how to mitigate for each one of the use cases. And the other thing is that we also need to learn more on data protection and data governance around models because the case is that when you work with, for example, your GPT license in your laptop and another person is working with another GPT license but for the same company, all of that information travels. That means that you upload something, a query, for example, and perhaps you also upload a document. If you are going to use a retrieval government's generation, you need to read that document to strengthen information. So once you are uploaded, the data travels where the server is because in that server is the model, the foundational model hosted. So the data needs to be encrypted when it is traveling, then when it arrives into the place where the server is, it needs to be encrypted again so that the model understands the data, it is processed and then it goes back to the user where the user is located. The problem is that, for example, with the European data space, you are not allowed to take data out of the European Union in, for example, a server that is in the US or in another country. So that makes things complicated to build models and to use perhaps some companies or solutions that do not have servers in the European Union. And also how the data is being transformed, will data remains a part of, we can say, learning data for the model, we also need to know that because if a document needs to be protected, you cannot have that data being part of the training data of the model. The model needs to operate and forget about the data it just received. And all of that because this field is very human, we are learning as we work. We see a problem, we go, we discuss, we adjust and then we repeat, we iterate. That is the process more or less. Yeah. You mentioned already Cody, what are your experiences with using generative AI for coding? How does it especially work out in our regulated family environment? The thing is, I don't know if I'm going to get in a lot of trouble for saying this, but we have two, I think, visions on this, on coding with generative AI. We have one vision that I think is very conservative one in the sense that coding needs to be restricted only for AI and data scientist engineers. So you come as a normal human being of the world with an idea, then you transmit the idea to them and then they be like code. And we have the other vision that I am a little more inclined into that vision that actually generative AI is making coding available for people that do not have coding backgrounds. That means, for example, I would like to develop a webpage that has a lot of coding in the back end. So I just write what I want, how I want the web page to look like. And then I will have that code be in generative. First of all, it helps a lot of people that want to go into coding to just do a small steps and to perhaps when you have a bug in the code that you are trying to learn and to produce to have a model explain to you why you are having that problem. That was something that, for example, I didn't have when I was learning to code and that event that sometimes for some functions that were, for example, you were working in a pipeline, right? And then for some reason, the environment updated the versions of some packages. For example, you had a small change in the way that the input was received or for example, you have this kind of, okay, in this first part of the function, it goes the dataset. Here it goes, which part of the dataset you want to analyze. And perhaps because of design, they just switch to this. So you would have first to choose which part of the dataset you would need to use and then the dataset. That thing. From one moment, the pipeline was working. From the next day, you just open your computer. You just said, okay, I'm going to continue working with this code. You just hit it run. I give you a back there and you spend years, we can say weeks, reading things on blogs, seeing other things related to that code until you see someone said, remember that when, for example, I don't know, will go up, change this new version of the environment. They changed this package for this one and the positions are inverted. Once you realize that, perhaps you waste it two weeks. Now you just put that function into the charge BT, for example, or Claude. And you can say, why is it not working? And then it's going to say, because this version of this function has the positions inverted. You just go and change that, no, even you can ask BT to, can you please give me the correct version of this code and is going to rewrite the code with the new fixed function? So then your problems go away. So in my case, I think that is, you wasted a lot of time in something that was so mundane of just a change in a position. And I think generative AI there is a very helpful ally or, for example, when you have the idea in the sense you are speaking with someone that wants to develop an algorithm for images screening for ABCs. And they are telling you, yes, okay, so we have four line images that we need to identify. So this image, we have the data sets with the labels and we would like, perhaps, a neural network, because that is some algorithm that works super nice with images. So you know in your mind, okay, the images are going to come here. The director is here, the data set with the labels is here. So I need to combine this and this. So if you need to sit and write it the all way, perhaps you would waste, I don't know, finding a repository that has a pipeline similar to the one that you want to develop, then write it on your code processor, then you know, testing. But now you can say, Chargingity, hi, here is my problem that I need to solve. I need to classify these four categories with an neural network. It could be this one or this one. Here is the format of my data set with the labels and here is the directories. Can you build me a pipeline for this? So then you have your entire pipeline in seconds and then you put that into your code processor, you hit run and it works. Then of course you made slight changes, perhaps when you see that curiosity that is not working, the way that you want, you could add augmentation techniques or themes to improve that. But you just have, we can say 65, 70% correct pipeline in 5 seconds. So there is the nicest things about this. You want, for example, to write or to do research in a specific area. You now can use, for example, OpenAI to do what is called a deep research. So you can say, Chargingity, can you please build me a landscape of these and these subjects? What is done? And where are the admin needs, for example? And then it's going to go and bring papers, blogs, any piece of information with the reference until you, from these subjects that you want to know, here and here are just two publications. That they are not addressing this and this problem. There you have your way to go for research, for example. That is super helpful, for example, when you write about introduction of a paper in these kind of areas. And I love that it comes now with the references. That was a big limitation in the past, said you didn't know where everything was coming from. And so it was Ross, really, love from it, from a transparency perspective and of course, you couldn't really use it for research. And now it's much easier to identify what is actually fake and made up by the model what was actually real data, real publications. I also think that when it's about coding, I know that many struggle with coding for figures on nice figures. You could spend a lot of fine fine tuning figures. And I think you do actually spend some time fine tuning figures, but of course with AI, it can be much faster to explore a couple of different ways how to approach the figure and you can get suggestions on improving figures and so on so that hopefully leads to much more use of data visualization since the future. Exactly. Actually, Alex, you hit one of the hot topics also for AI that is this kind of data quality control and data standardization when you have, for example, in some cases that we have this kind of migration that is happening from SAS towards R. So yeah, it's going to take a while because SAS is a language that has a lot of validation behind and a lot of years of use. So we practically know anything that can happen in SAS and R is just the opposite is something that is starting to be used. The thing is some things as you probably know in clinical trials are done in SAS and some things are done in R, right? And there is the moment of truth when you are building the reports and building the documents that you see perhaps a table that comes from SAS. And then a second table or a figure that comes from R, they have different data inside and it's the same study, it's the same dataset and the amount of human hours that it takes to go line by line and say, here it should be 15, why is 13 like where is this number coming from? And then you go and see and perhaps when the function was written in R, there was a skip in the line. So you're not seeing the same role you're seeing or will be behind those things are the use cases for AI also in the sense that you say here is this other figure table. Can you read both and tell me where the mistakes are? The language model is going to say here line, I don't know, two in this table should correspond with line two also in this table, but in this table is line three. So you need to mind the problem. So instead of taking, I don't know, perhaps sometimes 15 hours and this is something that is also very interesting that as humans, the more tired that we get when we are reading the same thing, the harder it's for us to support problems, AI, it doesn't get tired. So it just go over and if you say, can you please do the same analysis iterate 10 times, it iterates 10 times. So then you see, if there are some different answers, you can say, okay, from I don't know 10 attempts, 8, give the same response, it could probably be that the right thing. But imagine a human can you revise this table 10 times like, yeah, that's the main difference. So that is another case. And also building, for example, very mundane things that are that take a lot of time from people, the kind of trained people with BHDs and with, I don't know, 20 years of experience writing the food notes or writing captions for figures or for tables. That is the case for AI. Yeah, completely agree. Thanks so much for that great discussion. We touched about various aspects within biometrics with medical writings, statistics, coding, heavily talk about medical writing and coding, and how that can help us to become more effective. What has the different use cases there? Now you also talked about limitations that we need to take, have into account where we need to be a little bit careful in terms of uploading data, for example, there is lots of opportunities. But we need to also be aware about our space that we are working in and the specific constraints that come with this. Now if the listener says, my company would really benefit from all of that and we are not a huge company with lots of AI specialists already, how can that company benefit from you and your colleagues services? The first thing that I would say to them is that, of course, they should come to me, to us, sitel with any kind of questions that they would like to know, but the thing that I would say is that they just imagine what they would like in the sense, because sometimes when these fields of AI feeds a lot from abstraction and imagination, in the sense that when you know the process, you are the expert. In the sense that I am an expert in AI, but I'm not an expert in the process. That means that each company, each place has their own processes. For example, data governance processes, people that need to approve some data flows to move to the next stage. So each company is a world inside. And the more you know and the more you tell us about how the process goes, the better for us to say how we can pull AI into that process in a way that is organic and that it respects the data governance processes of the companies. I have seen a lot of very nice solutions in the sense, very well thought and with a lot of justification behind, as use cases, to fail because the processes inside the company that this solution was going to be applied on weren't totally known and were took for granted. In the sense, oh my god, it's going to be just the same as this other company, no big mistake. We need to know and we need to perhaps meet a couple of times where you tell us, yeah, we have this server. This server is located here, the data from these trials come you know into this way here and then the data needs to be approved by this and this and these people so that it is available for us. All of that makes us to say, okay, so perhaps the language model needs to be hosted here in need to have access to this and this data. We need a person to approve this process so that we also protect very carefully the data of the company so that you also have patients trusting that their data is protecting because one thing that is really damaging for AI is when the data is not protected because the solution could be amazing in the sense of predicting what you need to be predicted for example a disease risk or I don't know at time to an event that you want to prevent in a patient that is amazing but what happens if you know that algorithm is leaking data to the outside and a patient then says no, I'm not going to be using this algorithm because I don't trust that my data is going to be secure there that is a thing that we cannot afford because patients need to trust that the algorithm is going to take the best possible protection to their data and also that the output is going to be something useful for them so that they can use that and we can improve health systems with this. Everything needs to be needs to be considered and the other, the last thing that we're going to say about this is that for me especially this field is also filled with magic a little so you need to bring your happiness here in the sense what would you dream of having with AI what would you like this to look like at the end for example I would like this document to be completely automatic though with those things we can work in the sense okay let's see how the document is being produced perhaps we cannot produce it completely in the first iteration of the process but we can say okay let's do the synopsis. To once we validate that we continue with the rest and in no time you have the prototype there ready to be used but bring the thing that you want bring the thing that makes you happy like the sky's the limit with that kind of approach we can speak we can iterate we can if we don't have the answers for this just now we can have it for the future but we don't we need to come with our hopes because that is the only thing that is going to give us this strength to push forward when we want something to be built and change the life of people so that is basically what I would say to anyone but wants to do AI with that awesome awesome and we will link to Manuel Lincoln's profile and his email address on the homepage so just check out the affectives that's decision.com and then the podcast you can find all the episodes if you use for a little bit and search for Manuel in the costume thanks so much for that thank you so much our next out AI I have the feeling that was potentially not the last one so this show was created in association with PSI thanks to Rainy and 13FVS well put the show in the backhand and take you for the next free trip potential re-price room to the station just be an effective statistic.

Podcast Summary

Key Points:

  1. The guest, Manuel, transitioned from a medical science background to specializing in AI after encountering its potential in pharmaceutical projects, undertaking rigorous training in Barcelona.
  2. Generative AI is most effectively applied to repetitive, structured tasks in pharma, such as clinical report generation and medical writing, but faces challenges like hallucinations and data governance issues.
  3. AI tools significantly aid coding and data analysis by speeding up development, debugging, and research, while also enhancing data quality control and visualization in regulated environments.

Summary:

In this podcast episode, host Alexander Schacht interviews Manuel, an AI specialist with a background in molecular biology and medical affairs at Sanofi. Manuel explains his career shift to AI, driven by a need to engage in technical discussions and innovate patient care. He describes his intensive training in Barcelona, emphasizing hands-on coding to deeply understand algorithms.

Manuel highlights that generative AI is best suited for repetitive, structured tasks in pharmaceuticals, such as generating clinical reports and medical writing, where data formats are consistent. However, challenges include model hallucinations and strict data governance, particularly under EU regulations that restrict data transfer. The discussion also covers AI's role in coding, where it accelerates development, debugging, and research by providing instant code generation and analysis.

Additionally, AI aids in data quality control, such as comparing outputs from SAS and R, and improves data visualization efficiency. Overall, AI is seen as a tool to automate mundane tasks, allowing professionals to focus on higher-value work that requires human intelligence.

FAQs

It's a weekly podcast that helps listeners achieve high-quality science, serve patients effectively, and maintain a great work-life balance, featuring hosts Alexander Schacht and Benjamin Piscat.

You can access free resources and premium courses at the Effective Statistician Academy on zeeffectivestatistician.com, covering various topics in statistics.

AI is widely used for repetitive tasks like generating clinical reports and transforming data formats, where it automates processes to improve efficiency and reduce manual work.

Key challenges include managing hallucinations (incorrect outputs), ensuring data protection and governance, and complying with regulations like EU data residency requirements.

Generative AI helps by generating code pipelines, debugging errors, and explaining code issues, making coding more accessible to non-experts and saving time on mundane tasks.

AI is slowly being adopted in medical writing for document generation and in data quality control, such as standardizing outputs between SAS and R in clinical trials.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.