Go back

#754: Accelerating healthcare decisions with agents

35m 55s

#754: Accelerating healthcare decisions with agents

In this episode of the AWS podcast, host Jillian Ford interviews GGUN and Kenji Fujita from CoHear Health about applying AI agents in healthcare. CoHear Health focuses on reducing the 20-30% of U.S. healthcare spend wasted on administrative tasks between payers and providers. The company emphasizes trust and transparency, ensuring AI systems are reliable, hallucination-free, and clinically safe. Domain experts are involved from the design phase, and the team uses evaluation-driven development with strict metrics to quantify performance. For high-risk scenarios, like denying patient care, AI is never used without human oversight, while lower-risk tasks, such as approving diagnostic imaging, can be automated. To incorporate domain-specific data, the company prioritizes understanding the motivation (e.g., accuracy vs. context) and carefully manages data use rights. Technically, CoHear Health adopted AWS Bedrock Agent Core, citing its speed, safety, and ease of use. Key features like memory client and MCP gateway allowed them to quickly transition existing services and develop new agents with minimal operational overhead. AWS workshops and documentation were instrumental in enabling secure, scalable implementations, significantly increasing their development velocity. The episode highlights how regulated industries can leverage AI agents effectively by balancing automation with rigorous human oversight and domain expertise.

Transcription

6508 Words, 35744 Characters

English
This is Episode 754 of the AWS podcast released on April 7th, 2026. Welcome, everyone, to the AWS podcast. I am your host, Jillian Ford. And this episode today, I am super excited about. I think there's going to be something for everyone here. I know agents is really top of mind for, I mean, let's face it. It's like every single person on the planet is probably thinking about this right now. And you get to learn from two people who have been in the trenches at a company that has not only just been thinking about this, but actually has business critical applications that are using agents today. So I'm really excited to talk to GGUN and Kenji Fujita from CoHear Health. So GGUN is a chief data and AI officer. And Kenji Fujita is the staff AI platform engineer at CoHear Health. So there's something here for everyone, whether it is you or someone who is thinking about agents and how do you apply it? Maybe you're in a highly regulated industry. And you maybe you want to understand how to use it in AWS, some advice from these two. We're going to cover all of that. All right, let's get started. So GGUN, I'd love to understand first if you can tell our listeners about what is CoHear Health and the specific biggest business problems that you were thinking about within the healthcare industry. Oh, first of all, thank you for having us on your podcast, Julian, what a privilege. So CoHear Health, we are a clinical intelligence company. And our mission is to streamline the pay-and-povider connectivity and the collaboration. So just take it a second when I say pay-and-povider, what do I mean by that? Here's our insurance companies, our nonprofits, or government entities that finance healthcare services, where providers are entities like your doctor's office, a hospital system, provider groups who provide care. So I guess providers provide care and pay-and-povider. We have millions of providers in the United States and hundreds of pay-ers. So you can imagine with these two entities in order to support the whole healthcare ecosystem, the non-alonger transaction, a lot of administrative tasks, so we put pie-er-arth, claims processing, payment, quality, care coordination, and unfortunately the fault-wasting abuse that comes along the way when you have all these back and forth. And if you look at different studies, most recently, Health Affairs published an article where about 20 to 30% of the healthcare spend in the state are spent on administrative tasks, 20 to 30%. And somewhat them are necessary, somewhat them are avoidable, or I would say low value. And depends on which studies you read, it's about half of this 20 to 30% low value administrative tasks. So that's about half a trillion dollar. So that's the problem space for your health to solve. We want to eliminate the waste. When you do that, what does that mean, right? Patients can get the right care faster, providers can actually focus on what they do best, and the payer can really do a good job financing to care. So that's the national. Wow, there's just a lot here, I think, especially from healthcare, but I think other folks who are in different industries can even see some parallels to some of the challenges that they're thinking about. So when you were addressed with these challenges, how do you think about it in terms of implementing AI solutions? Yeah, it is a big problem, but it's also a very personal problem. Healthcare is very personal. Even when you're simply talking about getting a bill that you don't understand why, or having to wait a few weeks to get an answer for, to get an image, it's a very, very personal. So as we think about AI solutions, it really has to do with trust and transparency. It is the most important thing is the trust. And remember, I was just reading to my kids, Bernstein's books. Bernstein's fairs books. Love those books. Right. And this is book that talks about once trust is broken, you can get it back. So I think as we as technologists and think about building and rolling out these AI solutions, we have to get it right the first time, which is fair and forgiving in this notion of the cast and the deterministic world. So that's something that it's top of mind for us. So what does that mean? We need to make sure that what we do is reliable and consistent, there's no tolerance for hallucination. But no, also there are a lot of important experts opinion we have to take into account. It has to be clinically sad. A lot of literature, a lot of guidelines that we can count on. But the funny thing is when you put your doctor in this room, it's for the same case, chances are they don't agree on everything. That's why we love to get a second opinion, right? So as we think about building AI, we have to take all that into consideration. What do we mean when the system is performing? Based on whose opinion is that? And while the new one says that we have that account for, where we really need to say human needs to be in the loop, I think every last but not least is the security and privacy aspect. And we can get into somewhat the as we think through when we kind of get our hands dirty and building a system. That's a lot of technology safeguards, but also process safeguards that we have to consider. These are some themes that I think a lot of businesses are thinking about regardless of what industry they're in, especially with AI. I think the bar just keeps getting set higher and higher, which is great because now the customers want to be able to implement to be able to serve their end customers is going to be an even better, more accurate, more performant application for them. So I'd love to understand how are you thinking about no hallucinations, ensuring that it really is a safe and reliable application? Think let's pull it in terms of the software development life cycle. Yeah. So when we are doing design and kind of the reference architecture, it's easy to have the technologies in the room. I think, especially in healthcare, now imagine in many, many new ones, domains, we must have to expert in the room on a get go. It is common for many healthcare startups to talk about performances, talk about clinicians in the loop, but I think there is a difference when you engage your domain expert in the beginning versus at the end when we simply ask them to do validation. In Korea, we do every single development project have clinicians on the team and in the loop, as opposed to waiting until we have already built the prototype or already about the largest solution asking them to validate. And I think that domain expertise is key to ensure that we're measuring the way things. And I think that leads to my second point. I grew up in an era where we talked about test driven development. And I think now, especially with a lot of this agentic solution, we really have to move to the mindset of evaluation driven development, where e-vow comes first. What are the metrics that are important? How are we in a track? We talk about no tolerance for hallucination. It took us a month to iterate on exactly how we quantify hallucination in particular clinical settings. So all those often work is more important than ever. I think that's key. And once the part is where once the solution is launched, now we're at the monitoring and tracking phase. Obviously, I haven't tried to force that for a moment to turn this key. I potentially using the genAI itself to help us judge. But I do believe in the importance of human audits, having that regular, legitimately same-poly human audit. Really, really critical. And I'm going to say one last thing. And Kenji, you may have something to add to because we've been working together on this. It's relearning that there's different personas that will use AI solutions, especially since health case such a personal space. And even learning about how the year-old out to different persona groups over time help us build a more trustworthy and useful applications. Yeah, the one thing that I would add there is the key focus for us at a lower level is having strict guidelines and standards so that all patients get unified care while allowing for some level of user preference when it comes to how that how that care is received. There's so much to impact. So I'm glad that we've got more time. I'm going to ask so many different questions. But let's dive into that. What Kenji was just talking about that user, you'd use the words I think user standard care, but really like focusing on the end user. And that sounds like you also need to be able to have that domain-specific data in order to be able to bring it back to providing the best care or the best experience for that end user. So maybe you can tell us how companies can really think about incorporating their own domain-specific data in terms of AI. It's a great question. I think technically there are many ways. Why? Case-specific context or pre-training of models, the spectrum is wide. But I think instead of talking about the how I want to talk about the why and what, just for a moment. Ask ourselves why do we want to incorporate domain-specific data? Is it because I want the solutions to one faster, one cheaper, be more accurate, be more contextual? Is it a matter of that you want more co-op control over your system's output, especially given the industry operating and to Kenji's point where there's a standard of care that we want to maintain? I think understanding the motivation behind of using these domain-specific data will help us pick the right data and at what point do we incorporate them? So there are instances that will make sense to incorporate more of a knowledge system. Like in our world, it's more around medical society guidelines, standard of care, ontology framework, incorporating them will allow us to have a more consistent framework in how How to AI. I operate, but then when the goal is to provide more contextual accurate response in every single interaction, then we need to have our AI, biode access, faith, specific individual case notes. And they are well, and they are important. It goes back to, and last but not least, put it plug into the notion of data use rights isn't really critical. And when we entrust it with patient data and when we entrust it with business sensitive data, what rights does cohere health and our partners have to use the data for what purpose is really, really critical. And it's quite nuanced, right? Are you using it for operations versus are you using it for learning or are you using it for educational training or that nuances have to be a kind of for that? That is such a good call out because I know I've seen companies already start going down the route of maybe using a specific data set for an example, and then it's the engineers who get really excited about the problem, going to find out later on that like, oh, sorry, we can't use it because of like the rights to the data for maybe it's like compliance licensing whatever kinds of reasons. So that's a such a good call out that I think will definitely help a lot of the listeners. Same with what you were saying about having those experts really from the early stages of the process, instead of the human in the loop just being the end part. So maybe you can help us really share your thought process of how do you decide when to actually automate versus actually having that human oversight that's part of the process? It's the million dollar question, right? It's risk and reward. Human in the loop makes a lot of sense when we're working with high risk, high reward cases. And like for instance, in our case, we have a good number of solutions that target prior authorization automation, like trying to get the patient to the right care faster and with less paperwork. But we've made a clear decision regardless of regulatory and which geographic geography we operate, right? AI will never use to deny a patient's care or to even deny a provider's request to cohere health. We always draw a clear line on when we were not out of it is when we have to say no to a patient's care. Going back to your question earlier, Julian, about human oversight, even in the cases where we choose to automate, I think oversight is still essential. I think it's just a matter of at what point is it coming to play. Oversight should always happen with the system design, reference architecture and how we design the evil and oversight should always happen with on it. But I think it's in the high risk scenario where human oversight is in terms of every single case and every single nuance detail. If you're doing that middle ground, I think it's the key to figuring out a scale. Yeah. Maybe you can give us like an example that I think can help the listeners maybe visualize what that could kind of look like in their own business. Yeah, sure thing. So for instance, going back to the prior authorization automation example, we could look at a variety of clinical areas, right? Like with a lot of us have experience getting a prior offer imaging trying to figure what's going on, right? A diagnostic reason. Before you go get paid off because you're going to get a knee surgery, depending on the clinical use case, the risk tolerance is quite different. Like it's something when you need to bring me to your operating room and cut me open versus getting an MRI which is taking happening out of your day and with minimum radiation exposure. So really considering that the way it could help a project is for a knee surgery decision, we'd better be absolutely confident before we automatically say yes to the case without a human review. So as for a diagnostic imaging, we will likely say, as long as there's no control indications, as long as patient risk is considered, as long as it's covered by your policy. So there's no financial risk. We'll go ahead and say yes without a human intervention. So that new answer, it's, that's why the human experts are domain experts in the loop that's redesigned the system is so critical. I love that. So I'm really having a framework for with the experts assessing the actual risk and then using that to be able to design how the human in the loop human oversight is part of the entire process. Let's get into how this has actually been built. So Kenji, I'd love to understand really what are some of the factors that your team was thinking about that ultimately led you to choose a to Amazon bedrock and Amazon bedrock agent core. Sure. Yeah. So it, I think timing was a huge factor for co here. We this past year have invested a lot of time and resources into building out a platform around our AI and a lot of that incorporates how do you scale with the new agentic services that are that are out there in the market. And so we were attending workshops with AWS for for agent core. We were evaluating the different components. I think the key thing that stood out right away was the speed of innovation here, the ability to build these agents quickly, effectively and safely. A couple key components for us that I think most developers can get held up on our memory and MCP. And so it was, it was clear to me that these were paramount for the service teams at AWS when they were building out agent core because the tendency concerns that we have been a highly regulated space like healthcare are covered with some of the components of memory client out of the box and implementing them only takes a couple lines of code which to me is a huge, a huge benefit. The gateways is another thing. So we had started building out our own MCP servers, but with the identity built on top of gateway and the different targets that that the gateway provides, we've been able to at least iterate on our research and scale out the potential use cases for for our agents because again, it only takes a couple lines of code to implement an entire MCP MCP server target. And that's to be it says a lot that you're saying you're able to build it quickly, effectively and safely because better agent core is relatively new. So it sounds like there must have been some folks at AWS that really helped you. So maybe you can share some of some how AWS was really able to help you with everything you were saying earlier, like the technical challenges, the business challenges that you had and to be able to actually build it in production today. Yeah, I think what helped us was a little bit of hand holding around understanding our use case, right? We are top concern as always the security of the data when it comes to implementing a solution like this. And so the first thing we brought to them was how do we transition our short-term memory to agent core so that we have tendency separation? The workshops that we went through with some of the Jupyter notebooks that they had available had everything that we needed right out of the box. And so going through some hands on experience in a test environment was super helpful. And then understanding the documentation was also key. And I think the namespaces that the agent core memory client provides are very intuitive to the setup. I'd love to know based on your experience between the workshops, the documentation, is there anything else that kind of stood out to you that can help folks who are on that journey of implementing agents? Yeah, so I touched on it a little bit, but I think the amount of experimental research that we've been able to do has far exceeded what we thought we would be able to do by this point. And a lot of that is due to the ease of development here. So I'm trying to think of a good example to share. There are two different use cases that we have. One was transitioning an existing service over to agent core. So we had spent all of this time setting up and evaluating and setting up memory for a chat agent and transitioning over to the memory client was a relatively trivial task for our team to implement. And a lot of that was due to the documentation and the workshops that they were able to attend. But then there's also net new development. What's clear to me is that agent core really takes away a lot of the ops concerns from the MLE. So the developer on the ML side who typically wouldn't be a DevOps expert doesn't have to evaluate those concerns as heavily because a lot of it is handled by the agent core service. Back to something you were saying earlier, because I think this will resonate with a lot of folks who are listening. MCEPs are really a huge hot topic right now and a lot of people are thinking about building it themselves. So I'm curious from your experience when you were at that stage of you had started building it yourself and then you would start then use bedrock agent core. If there was any other learnings that you had from that experience that you can help someone else who's really thinking about building it myself or should I use bedrock agent core to make that easier? Sure. So here has been around since before a lot of this technology existed and so some of our applications have tendency built into them that the agents need to follow. So we want to make sure that the agents are following the same patterns that we already had in place pre-agentic implementation. And so we were building out our MCP servers and ensuring that we had off proxies to send through the same sort of approach that we would follow on the core application side. But the agent core gateway handles this implicitly with the identity provider. So it's one of those things that with a couple lines of code like I was saying you can just pass through the same authentication method that we would typically set up on our own with the configuration enabled by agent core gateway. So I'm curious now that you went from before you started to really like go down a path building it yourself then started with agent core. Did that change at all? be like your time frame of when it was a that you were able to put into production or any other areas of your velocity. - Oh, 100%. Yeah, we were talking about this all week. I think going into this quarter, we had maybe an agent or two that we had planned to develop and looking ahead at Q1 and the rest of 2026, it's full of agents, it's full of agentic systems, multi agents, and I think GG touched on this a lot, but the evaluations come first. So a lot of our focus has been on building out the evaluation framework this quarter so that we can scale and continue to build new agents that we know are gonna be successful on the first past, using H-A-Core. - I've got a few questions that I wanna get your both of your opinions on. So GG, I'll start this one with you. So model choice, this is a super hot topic that I know a lot of businesses are really thinking about. So I'd love to hear how you think about looking at evaluating different models and the term I love that you chose earlier, evaluation driven development based on your experience, maybe some insight that you can share on the business value of that approach. - Yeah, actually it could be helpful for me to go back a few years of history. Canji talked about how cohe health has started a few years ago before this agentic revolution. And in fact, we were using our own transformer models to do NLP months, if not quarters, before Chachibitikimat. So the company was already a solar industry on adopting this cutting edge tech. So I think we have a unique perspective 'cause it's always the question, feel versus buy or do we tune? So I think the decisions we should do we continue our own journey in building our own model from scratch versus adopting one of the frontier model with pump tuning or maybe goes on and a half of fine tuning slash maybe pre-training. So we've been having this internal healthy debates for quite a few months. So I think it's a unique experience that I would love to share more widely. And I think one thing is changes constant and the only, the best thing I could do as a leader for these amazing technologists is to separate clear metric, accuracy, cost, latency, reliability. And honestly, how much e-ball data is needed for each approach? We cannot push any AI out without publishing e-ball data, especially in the industry we end. So really thinking for all those metrics and I can be honest with you, Julian, earlier this year is still needing very heavy tourally self-training and more recently is leaning more and more towards fine tuning or perhaps using one of the frontier model at least at the get-go so that we can get really good coverage. So that I think it's keeping an open mind. What has really helped us is once we agreed on these are the business and operational metrics that are important to us and setting up a leader board. So that we have the ability behind a scene, not as part of the constant spring planning that we have a way to keep monitoring and tracking which models are winning. But what makes sense was to make the switch. So on top of an automated LLM as a judge approach, which we have in place, we also have clinicians labeling the data for us behind the scenes. 200% Kenji. And changes a constant, having a leader board to keep watching against metrics are so helpful. I think going forward is being really cognizant on how we collect ground truth data so that we can have that use case specific insights to make these decisions. - There is a lot to really unpack there that I think every single listener can take, get a takeaway from. Metrics, having a leader board, I know a lot of businesses out there that I speak to don't have a model evaluation framework. They're usually just sticking with one large language model and they stick with it until maybe there's a reason not to. But I love that you're really assessing all these different options that are out there, I think. And obviously it's very clear that you're looking because you have all these different metrics that you've defined ahead of time, you're able to then, and you've got this process, you're able to have the best cost possible at the lowest latency, at the best performance, which at the end of the day, that's what companies are all looking at. They are just don't have a system that can be able to help them get all of the benefits that they're really looking for. So I would love to hear your advice for a company that right now they're using a single large language model and they're curious of, there's probably other models that there are definitely other models that are out there. But how do they go from one to being able to assess others so they can pick one or more that are going to be best for their use cases? I think you summarized it well, right? Let's make sure you know what's important to you, one of your metrics, and then investing in that all I made at Eval Framework, both human in the loop and using large language models so that these decisions can be made with data driven decisions, and that's a big part. And when it does show that there may be value to switch, my personal experience says that it's not always one size fits on. It's not that you switch from one frontier model to another for all your use cases, or even within the use cases every single piece of your pipeline. So it goes back to architectural conversation I working with the architect and thinking through how you design an architecture that allows you to have that compulsibility, the ability to, for some, in our case, certain use cases rely more heavily on smaller models that we host internally in certain use cases rely more heavily on frontier models. But if your architecture doesn't support it and every time it's any built, then you make the adoption cost very unbearable. - Yeah, Gigi touched on it. I think having the platform, investing in the platform is key to this style of development. Both the Evaluations Framework and the Gateway. I think Agent Core Gateway is meant to be a gateway for the agents. I think having an LLM gateway or an AI gateway is also valuable to put on top of this framework so that you can iterate quickly, change targets, and evaluate these at a much higher pace. - So I'm very curious about, like, really the business impact that you've been able to see in AI. I know there's still companies that even struggle to be able to measure the ROI of AI. And so hearing, I think, from your experience, we'll certainly be able to inspire them. I know in addition to everything else you said earlier that definitely has, if not already. - Yeah, although, let me go back to the example I started earlier in this conversation. In the pile of space, we work with clients who have to deploy dozens and dozens of nurses and MDs and doctors to review these cases in order to just meet the volume and meet the turnaround time requirement. It's a highly regulated industry. We have 14 days to decide, but starting in a new year, you only have seven days to decide. So it's easy to just try to fill bodies in the problem. But with the AI system that we build out with our clinician in the loop, we are able to achieve 85% automation. So 85% of decisions are made within minutes. And for those 15% of cases that require high touch human review with our agent system, we've seen about 30 to 40% improvement in productivity. And something that's tough to measure, but we are hearing feedback from the users is, "Mix it with the helps with the job satisfaction." Because they spend their time making clinical decisions as opposed to trying to figure out where the information is or trying to dig through other requirements, everything is serviced, all the deep research is done on their behalf. And they can really just focus on the area expertise. So yes, we save time, but also, I think we have better retention. - That is definitely a testament to the operations that your team has done. I mean, to be able to get to 85% automation, your end customers being even happier with the solution, that really speaks to, I think, for all the listeners who are thinking about how to be able to get there. And especially those, I know who are really in a monolithic type of architecture right now, where they have one LM, and they're not able to maybe experiment with a number of different models that are out there to be able to maybe have a certain use case that's for the look, they can get it away with like a lower cost, one that only have them better latency, all those different factors that you were talking about. So I think I just love those metrics because I think it just shows the others who are in the early stages of their journey, what's possible. Okay, so other areas that I think I'd love for your opinions on. All right, so Gigi, based on what your experience, what are some, like, what are three things that all companies should really be doing when they're in the early stages of being able to build agents before they actually push into production? I think I'm gonna know your answer, but maybe I'll be through that. (laughing) - In a little context, right? I've been doing this line of work for 20 years. Agents are not, right? The movement from a proof of concept or a pilot to a large scale solutions. Even McKinsey says that 95% of AI prototypes and POC don't even go into production and only have one single production as a state production. So this has been a age-old problem, regardless of age and don't know. I think agent does create actually a pressure because on one hand, you can innovate and experiment faster, but on the other hand, there's more unknown whether you have to manage. So you actually exemplify the challenge. And you're right, we're gonna start with evaluation driven development. If there's one thing you're gonna take home from this podcast, there's the one line that I will really encourage us. And as we think about these metrics, to McKinsey's point, having the right experts to label, provide a label data, provide a grant to you, make sure that they tie back to your business and operational metrics. So they're not pure functional, non-functional evil, but tie him back to the overall company of your clients. The other piece I've personally seen a lot of struggles going from POC to scale is not having that not investing that time. Let me put it the other way. I should put it in the positive way. Let me try again. Now the part I've seen successes in taking from POC to large scale deployment is taking the time to define and letting everyone know what must be true for the agent to be successful at scale because the nature of POC is to simplify, right? It's to not consider certain edge cases and it's to assume certain level of integration and operation or efficiencies. So in order to kind of flip that switch, we must be very, very clear on what are the important criteria for it to be successful. And they're not usually AI related. They're usually about the people who are going to be using the tool. Now you're going to give them the right training, you're going to give them the right transition plan, are you going to bring in the right advocates because as I mentioned earlier, everyone look at it, technology differently, different personas, right? What do you bring on board to help you event? And oftentimes things fail because of processes. Right? Because if you don't change your processes, but you get the new tech, you're not going to see the benefits and you can quickly fold by the solution can quickly be minimized. And then I think the third thing is, it's like classic right system deal, right? We always assume data integration is easy and it's never easy. So I think that's the piece that we always have to kind of take a step back and say, amazing AI system. Let's make sure the people, the process and the data already and having that clarity so that everyone's on the same page and margin to a single. I think you just saved people a lot of time on that because people get so excited about AI and already start thinking about like what LM, for example, are we going to start using but getting the domain experts really in part of the process, making sure you really understand like the that you were talking about earlier, like that what your building is going to change the entire process for your end customers and really having clarity on what that means for them, are they going to even use it? Even the data ingestion part as well. I think these are all prerequisites that people really need to think about. I've seen that as well as they often become overlooked and then you have to take two steps back before you can go ahead. Anything you would do differently if you were starting over today. I think the Jerry is still out. We don't know yet. One thing that I'm still thinking through is it's more effective to have a small team that focus on the agent development in a larger org or is it better to plug the agent development into everything or development team? Let me try to say it the other way is that better to have a centralized agent development team, agent development team and let them drive the innovation and then dissimulate what they learn is that more effective or is it more effective to have each development team to start adopting and doing more of a federated model? Jerry is still out on that one. That sounds like that will be a part two episode. We will let you know, Jolani. I don't know. Can anything you want to say? No, that's a good point. I don't have anything there. It's a tricky problem. All right. I've got one last question for each of you. Can you start with you? So we've got listeners here that are in all different industries and all different stages of their journey within AWS. So for those who are considering building AI agents, what's one piece of advice that you have for them to get started? I think AWS has so many resources out there right now to make itself serviceable. So I would, my best advice is to just start testing out and trying the tools. And at least in my experience, asking the questions early and often to the service teams to our account managers, it's helped uncover the solutions that we would have spent more time trying to figure out ourselves. Can she is right being able to have hands-on experience is the best thing to make good decisions. And I guess one more thing I'll add is don't let formal get in the way just because everyone seems to be deploying and benefiting from agent systems. Really always start with the why you'll be surprised. There's still a class of problems that may not be agentic, right? And there's a class of solution that can be very well solved with a single LLM and just really going back to the business and success metrics to make decisions is key. I think my my advice to the folks who are wanting to dig into agent and so forth is think hard about the team you assemble to make this real. I think agentic work truly does require a new profile of developers and a development team. We mentioned earlier about the importance of having domain experts in the one day one. Those who are technology savvy domain experts, they are gold in the team. And we've had amazing success seeing the whole platform developer work inside by side with a data scientist and work inside by side with a clinical MD and a nurse. That combo really allows us to iterate super quickly and doesn't know framework. I don't think we've had that kind of dependency in terms of diversity and skill sets being in the same room at the same time before the agent. Wow, this was seriously a master class on I mean evaluation driven development AI driven development. There's clearly was something here for everyone. Thank you so much. Can Gigi, this was phenomenal. Really appreciate both of you spending time here with me on the AWS podcast. Thank you for having us. Thanks, Jillian. This was awesome.

Podcast Summary

Key Points:

  1. CoHear Health is a clinical intelligence company that aims to streamline administrative tasks between payers (insurers) and providers (doctors, hospitals) in healthcare, targeting the 20-30% of U.S. healthcare spend wasted on low-value administrative work.
  2. Trust and transparency are critical for AI in healthcare; the company prioritizes reliability, no tolerance for hallucinations, clinical safety, and human-in-the-loop oversight, especially for high-risk decisions like denying patient care.
  3. Domain experts (clinicians) are involved from the start of development, and the company uses evaluation-driven development with strict metrics to quantify outcomes like hallucination, plus ongoing human audits.
  4. Incorporating domain-specific data (e.g., medical guidelines, individual case notes) is guided by clear motivations (accuracy, context, control) and careful consideration of data use rights and privacy.
  5. The company chose AWS Bedrock Agent Core for its speed, safety, and ease of use, particularly for memory management and MCP (Model Context Protocol) integration, which reduced development time and operational concerns.
  6. AWS workshops and documentation helped CoHear Health quickly transition existing services and build new agentic applications, with features like memory client and gateway enabling secure, scalable implementations.

Summary:

In this episode of the AWS podcast, host Jillian Ford interviews GGUN and Kenji Fujita from CoHear Health about applying AI agents in healthcare. S. healthcare spend wasted on administrative tasks between payers and providers.

The company emphasizes trust and transparency, ensuring AI systems are reliable, hallucination-free, and clinically safe. Domain experts are involved from the design phase, and the team uses evaluation-driven development with strict metrics to quantify performance. For high-risk scenarios, like denying patient care, AI is never used without human oversight, while lower-risk tasks, such as approving diagnostic imaging, can be automated.

, accuracy vs. context) and carefully manages data use rights. Technically, CoHear Health adopted AWS Bedrock Agent Core, citing its speed, safety, and ease of use.

Key features like memory client and MCP gateway allowed them to quickly transition existing services and develop new agents with minimal operational overhead. AWS workshops and documentation were instrumental in enabling secure, scalable implementations, significantly increasing their development velocity. The episode highlights how regulated industries can leverage AI agents effectively by balancing automation with rigorous human oversight and domain expertise.

FAQs

CoHear Health is a clinical intelligence company that streamlines connectivity and collaboration between payers (insurance companies) and providers (doctor's offices, hospitals). It aims to eliminate wasteful administrative tasks that account for 20-30% of U.S. healthcare spending.

They use an evaluation-driven development approach, involving clinicians from the start of every project to define metrics like hallucination quantification. They also rely on domain-specific data, human audits, and strict guidelines to maintain reliability and consistency.

Domain-specific data ensures solutions are faster, cheaper, more accurate, and contextual. It helps maintain consistent standards of care and allows for personalized, clinically safe responses by accessing specific case notes and medical guidelines.

They assess risk and reward: for high-risk cases like denying patient care, AI is never used, and human oversight is applied to every detail. For lower-risk tasks like diagnostic imaging prior authorization, automation is used, but human oversight remains in system design and evaluation.

They chose it for speed of innovation, ease of building agents quickly and safely, and built-in features like memory and MCP support. It reduced operational concerns for ML engineers and allowed rapid transition of existing services to agent-based solutions.

AWS provided hands-on workshops with Jupyter notebooks and intuitive documentation, enabling CoHear to easily transition short-term memory to Agent Core and set up MCP servers with just a few lines of code, ensuring tenant separation and security.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.