Go back

Being AI Native at ServiceNow

24m 8s

Being AI Native at ServiceNow

This episode explores what it means to be "AI-native," emphasizing that it is not merely adding AI features but building with AI as the foundational operating system. Dai Li distinguishes between responsible AI (governance and accountability), ethical AI (philosophical values), and human-centered AI (designing for human interaction and agency). Dr. Elena Beaver introduces ServiceNow’s first-of-its-kind open-source AI model accessibility checker, which allows anyone to test AI models for accessibility compliance, promoting industry-wide responsibility. A cognitive science framework from a research leader categorizes human-AI interaction into use, disuse, misuse, and abuse, urging design for proper use to avoid over-reliance or rejection. The voice AI team shares their open-source evaluation framework for voice agents, addressing challenges like transcription errors in cascade systems (speech-to-text, LLM, text-to-speech) versus audio-native models. Engineers Ian Thurlow and Andrew Yen describe how AI tools like Cloud Code accelerate their daily work, acting as a senior engineer for code reviews and problem-solving, without replacing human expertise. Overall, the episode underscores that being AI-native requires embedding responsibility, rigorous evaluation, and human-centered design from the start, ensuring AI amplifies human cognition and agency while maintaining transparency and accountability.

Transcription

3755 Words, 21213 Characters

English
[MUSIC] Hey, everyone. Welcome to another episode of Service Now Insights. I'm your host, Bobby Bro. So the big question we're always asking ourselves or being asked more often is, what does it mean to be AI needed? For sure, it's not just using AI. It's not just adding an AI feature to something that already existed, but building with AI at the core. Thinking with it, designing with it, even working alongside it. But what does that look like from a philosophical standpoint? From a responsibility standpoint, and from the ground level of the people, actually building and using these systems? So for this episode, I've pulled together some great tidbits from some of our past episodes. And we've got Dai Li, who is our AI ethicist and human-centered AI strategist, who's going to give us the clearest definition I've heard on what responsible, ethical, and human-centered AI actually means, and why all that matters. And then we're going to hear from Dr. Elena Beaver, our global head of accessibility, who's going to share something genuinely surprising, a first of its kind tool Service Now built, to make sure AI itself is being used responsibly. And a non-Aronathon, a research leader who is going to give us a framework from cognitive science that will completely change how you think about AI use and even misuse. And then from the voice AI team, we'll hear from Tara Bogevalli and Katrina Stankowicz, who are building evaluation infrastructure for voice agents and giving that away to the whole industry. And finally, from Ian Thurlow and Andrew Yen from Engineering, who are going to tell us what being AI native looks and feels like from where they sit, in the code every day. So let's get into it first with Dai Li. So for context, we live in a very nuanced era, and any of these terms could be an entire course on their own. But at least in the line of my work, when we talk about responsible AI, ethical AI and human centered AI, here are the things that differentiate them. First, responsible AI is this umbrella term that encompasses the industry of research, policies and people that are working towards a safer type of AI. But in practice, it's the organizational and governance layer. It's about processes, accountability structures, risk management, compliance. It's the how do we operationalize responsibility question? When a company says we practice responsible AI, they should mean they have governance frameworks, oversight mechanisms, frameworks and guidelines, and accountability chains in place. They should have systems and structures that make good intentions actionable. The ServiceNow Responsible AI page, which has a wealth of resources, including white papers, our mission statements, our principles, and the responsible and human centered AI guidelines, that I was so fortunate to have the opportunity to lead the creation of with brilliant minds. But jumping out like ethical AI, what are ethics when it comes to AI? And ethical AI is more the philosophical approach to the research. It's the values and principles layer. It's the what should we do at why type of questioning, fairness, transparency, avoiding harm. These are ethical commitments. And ethical AI asks whether what we're building aligns with society standards, human values, and the bar that we've set for ourselves and for society. I can't emphasize enough that when we're talking about like responsible AI, it's operationalizing the conversations and musings of ethical AI. So in a way, ethical AI is the moral and philosophical backbone of all of these efforts. And then we get to my favorite because my background is in UX design and research, human centered day, human centered AI. And human centered AI takes into account the philosophical questioning and research, the governance practices, the operational layers. And then it starts to ask questions about how do humans interact and experience and in what ways are they impacted by the technology? It's about asking questions like, well, what are humans good at? What are they not so good at? What do they value? Does this experience preserve human agency? Is what we're building, enabling and amplifying human cognition? And that's the thing that makes the third layer, human centered AI. But does it actually work for people? That is a core design challenge for designing for, especially AI native experiences. Think of AI native experiences like instead of just having an AI product or feature to whatever you're working on, having it be like the operating system. So with that, we're one of the biggest design challenges we're up ahead. It, one of the biggest design challenges that we're heading up against is how to, when you have this ambient technology that is probabilistic, how do you empower the human further to contest, make informed decisions and navigate this system, especially because ultimately, humans will be held responsible for the work that they're doing with these tools and systems. It's hard to think about in a technical term or in a technology term. Are we creating discernment or are we teaching discernment? How we should approach these design experiences is treating the human like a subject matter expert and given their knowledge of their domain when using AI or when in an AI native experience. Can their expertise be leveraged to catch flaws, to know how to use the system in a way that's meaningful for them? Do they understand the limitations? Can we show them the limitations? Can they help improve the system over time with their expertise and in turn can the system further benefit them? So it becomes a symbiotic relationship that I shouldn't say is a standard yet, but is a bar that we're trying to hold ourself to. AI native as the operating system, not a feature. And we keep coming back to that because it reframes the whole conversation. It's not about what AI you add on top. It's about whether AI is baked into how you think from the start. And if you're going to do that, if AI is going to be embedded in how you work and what you build, then responsibility can't be an afterthought either. Dr. Elena Beaver, Champions Accessibility at ServiceNow, and she's going to tell us about something that kind of stopped me when we first heard it. A tool they built to make sure AI models themselves are built responsibly. ServiceNow in partnership with the Global Accessibility Awareness Day Foundation has developed the world's first, world's first drum roll please, AI model accessibility checker. So you have accessibility checkers that are these digital tools, right? And you can run an accessibility checker on any website, anybody can't. There are free versions, they're called things like wave and acts. And you know, you can run a checker and it'll let you know how accessible and interfaces. Now with the advent of everybody going crazy over AI, there's been a real concern around ethical AI. And that's it's all whole on topic. And we want to ensure that AI is being trained on correct information and that it's aware of accessibility standards. And so what we have done is we've created a way for anyone who's using an AI model to run it through a free AI model accessibility checker that we host on GitHub. And it's completely free. We just want to open source, we want people to take it and steal it and use it and run their own accessibility, you know, quirks on their AI that they're creating, especially as they're more and more proprietary LLMs that are coming into play. And so we would love at ServiceNow for you to use our LLM and to be engaging with us on the now platform on the power of our AI that is homegrown. We are really proud of it. We think it does a really good job. But if you are using other more established models, called the amac amac.ai. That's a web address that will take you to the leaderboard where you can publicly see how some of these very commonly used AI models score essentially. And if you are going to just out of the box, pick an AI model, might as well pick one that scores well, right? So I'll try to be aware that different models are better or worse, added hearing to and reminding users of accessibility standards. And so if we have the ability to be using AI for good, which we do, let's ensure that we're using it excessively. Open source and free. So please go test your AI model right now. And that's not a company protecting a proprietary advantage. That's a company saying, we want the whole industry to build this way, which brings me to framework. And a lot of people think that's not really where AI is. It's not in the framework. But actually, if you're building it, managing it, or using it every day, you really want to think about the cognitive aspect of it. So let's hear from a non-theiranatham. Using automation or AI, you could interchangeably use that. Means designing systems appropriately so that the human and the automation or the AI can effectively and efficiently get the job done. That is proper use of that. Now, you think about disuse. Disuse is when the system is designed in a way, poorly designed, that the human is not able to understand what is happening. And when that happens, the human just shuts that off and turns it off. So think about an LLM that's not accurate enough in this era. And then comes up with articles. And if you search for something that comes up with the results that are an accurate, this is not good. I'm not going to use it. That is what we call disuse. Then comes misuse. Misuse is when you completely rely over rely on the technology be it automation or AI. And that happens when you have over trust in a system. And when you have over trust in a system, complicancy sets in. And when that happens, if something goes out of whack to my earlier point, everything is a problem. So in that case, in a misuse case, you really-- the human has no idea what's happening because the human is so complicit. So you don't want that either. An abusive automation or AI is when you just completely-- you just take that technology and just because you just apply it across every single thing. And that is what's called abuse of a technology be it automation or AI. So if you can really think about in the context of AI is use and designing for proper use, that's when you get the most success. And you want to stay away from misuse, disuse, or abuse. Hopefully that gives you a little bit of a framework to think about how to design for the best man machine systems. So use, disuse, misuse, and abuse. Now that's not to villainize AI, quite the contrary. It puts the responsibility exactly where it belongs, and how we design it, and how we choose to engage with it. So let's take a look at how AI actually works in practice here at ServiceNow, starting with one of the most interesting research projects that we've had on the show. The voice AI team here has built something remarkable, a rigorous evaluation framework for voice agents, and then open-source it so the entire industry can use it. That's not engineering. That's native AI thinking. So let's hear from Tara Bobo-Valley and Katrina Stankowicz. Yeah, so we're super excited about this. We're trying to release this open-source evaluation framework like you mentioned to basically be able to show how good our voice agents. What are they bad at? What are they good at and where are the holes that need to address in the future to make them even better in production systems? So what we're working on is a framework where basically you have some data. In this case, we have a domain that's focused on flight, like airline scenarios, like flight re-booking, cancellations, stuff like that. We're going to have this framework that you can use to generate these simulated conversations. And then we're also going to have these metrics that we discussed that will tell us based on the conversation, like how good was this conversation on the two dimensions that Gabrielle mentioned, accuracy and experience. And so we're hoping this will help basically set the standard for what voice agent evaluation should look like not just at service now, but like industry-wide. So when you're talking to a voice agent, there's usually two main architectures that are going on behind the scenes. And a lot of people don't know that often we have a cascade system, which means that there's actually three models working to try to understand what you're asking. So we have a speech-to-text model, which we call STT usually. So this is going to interpret what you're saying and translate it to text. So we need to translate it to text because often the most intelligent models, the LLMs, can only understand text, and they can't understand any audio. So when you say something, it goes through a speech-to-text model, translates it to speech, and it goes to your LLM. These are the LLMs that we generally see and use online when we're talking to chatbots. This LLM will understand the text that was given to it, try to accomplish its task, maybe ask you a clarifying question, and then output some text. Now we have a third model, which will take this text and translate it to speech to speak it out back to you. So these three models are working together to give the illusion that you have a voice model speaking to you. The other option that we have are audio native models. So these are speech-to-speech models that can actually understand and speak audio back out to you. The main trade-off though is that this speech-to-speech model is combining the work of three separate models all into one. So we lose a lot of the intelligence. So what we gain in experience where it's a much smoother kind of transition and it's much quicker to respond, you'll find that it's not able to accomplish tasks, the way that the cascade models are able to accomplish it. The drawback of a cascade system is-- one of them is latency. So you're talking to three different models, and it takes some time to pass the information back and forth between those three models. But then also when you are transcribing some audio to text, you lose a lot of the like emotion-handling or tone that maybe a user is trying to convey. So this LLM now only has a text version of what you're trying to say and can't really tell if you were angry or maybe there's some urgency in what you're trying to say. None of that will get translated at all to this LLM. Yeah, that's got to be one of the big issues, especially when it's a chatbot, because it's based on customer service first, right? So I'm angry. I have a problem. Fix it now instead of I'm angry. I have a problem. Fix it now, which has got to be a very strange way of interpretation, which I know we're going to get into. But there's some other things we want to talk about first, Katrina. And that's the problems of the initial translation would you put it like, what are some typical goal examples where we might have this cascade failure, so to speak, when it happens in voice agents? Yeah, exactly. So that's often one of the first points of failure where this transcription does not go as a plan or as you'd want it to. So for example, if you are trying to say your name is Bobby Brill, maybe it gets misinterpreted as Bobby Grill or something like that, or maybe Bobby Brill-- Happens all the time. Happens all the time. Yes, yes. I answer to Mr. Grill quite often. But the problem is that when this LLM receives the text, it kind of receives it as fact. So it doesn't really realize that maybe there was some misinterpretation going on there. And it's trying to accomplish its task, thinking that it got the correct information. So if you're trying to maybe find a reservation that was under your name and the LLM interprets your name as something different than what it was, then it's going to wrongly say that your reservation doesn't exist, or it can't find this information. And you have no way of knowing that it was using the wrong information versus your reservation, not actually existing. Like you don't have that feedback, or you don't always have that feedback of, did it use the correct word? And then on top of that, there are names like Brittany that can have many different spellings. And then in that case, again, will the speech-to-text model actually interpret the spelling of your name correctly and pass that information onto the LLM correctly? That's very, very common for it to make a mistake. And then it kind of disrupts the whole flow. And it really can't accomplish that task when it's using the wrong information. Yes, Bobby Grill, the Grill Master, summer's coming. So maybe that's a totally different podcast. But what you just heard is what makes AI native engineering different from just building something that seems to work. The voice AI team isn't just guessing at whether AI agents perform well. They've built the system to measure it rigorously. And then they've shared it. That mindset, build it right, measure it honestly, make it available. That's what being AI native looks like at the company level. Now let's get into the engineering. Because of all this philosophy and all the frameworks and team deployments, there's also just the daily reality of people building software here at ServiceNow with AI as a constant presence their workflow. Ian Thurlow and Andrew are going to tell you exactly what it feels like to be living in the code every day. In some ways, things are completely different, but in other ways, they're basically the same. We're still the same team working on new features and existing defects, but we've been sort of handed this really powerful tool that essentially changes the way that we execute that work. Lately, I've been leveraging Cloud Code to ask questions and respect about the code. Someone asked me a question about the code, and maybe it's some code I haven't looked at in a while. So I say, "Hey, Cloud, how exactly does this work? Show me all the classes that matter." And I start digging into that. So that's really helpful for me. I've also been using it to assist with code reviews. Say, "Hey, Cloud, I have this code review from one of my engineers. Review it for me." And sometimes it helps. It does catch some stuff that I wouldn't have probably caught on my own, because it really looks at the code and tries to run it, and it can figure out possible issues there. The one piece that has changed is, I really start with AI when there's a problem in front of me. Because you used to start with Google, right? But now you're starting with Cloud, or you're starting with HGBT or whatever. No matter what the problem is, it could be like, my dishwasher is broken. The first thing I'm going to do is say, "Hey, Cloud-wise, my dishwasher is broken." And I think that's changed for almost everyone in the world. We've gone from this, digging through Stack Overflow, digging through Java Documentation to having all that in one place for us. And that's really the. So it's really an accelerator more than really changing the world. It's making us do our work better. I can have a conversation with Cloud, and it really helps me dive into my code. It's as if I'm speaking to a senior engineer who's pretty well-molged on the code at hand. So it makes it a lot easier to do independent work in that sense. And you can't just dump your brain into Cloud, which would be pretty cool. I do think that the thing that always. that people always bring up is the calculator was going to make mathematicians obsolete. But that's not true at all. It made it so now they don't have to do these basic or even complex things on paper. Now they can do it in an automated way. But we still have people with mathematicians, with those statisticians, with all these people that are applying those skills. And I don't see how this could be different. And I do think there are some jobs that will be more affected and will be more accelerated by AI. And it'll be interesting to see how that sort of comes to fruition in the next few years. To piggyback off the ins comment, like the fundamentals are really important now. Because they're not really gated again by your development speed, or like if you know how to use a certain framework or a tool or language, for example, it's like, if you know the fundamentals, that's like a solid bedrock where you can start building anything also. And like now it's like, oh, you could prompt cloud, for example, to build this tool in like this language. But you still need to know the basics of like how pieces interact and stuff. And so having that as a basis is really important. And I feel like that's more so than ever now. So there you have it, fundamentals first. Because it's not about chasing the newest tool or applying AI to everything you can find. It's about having a strong foundation that when AI accelerates things, it accelerates the right things. That's what being AI-native means. It's not just a product or a pitch. It's a way of thinking that runs through every layer. So check out the show notes for links to all the original five episodes. And of course, head over to servicenow.com/docs to get insights into how to make this platform work just for you. I'm your host, Bobby Brill. Thanks for listening. [Music]

Podcast Summary

Key Points:

  1. Being AI-native means building with AI as the core operating system, not just adding AI as a feature.
  2. Responsible AI is an organizational governance layer (processes, accountability, compliance), ethical AI is the philosophical values layer (fairness, transparency), and human-centered AI focuses on how humans interact with and are impacted by the technology.
  3. ServiceNow developed the world's first open-source AI model accessibility checker to ensure AI models are trained on correct accessibility standards.
  4. The use-disuse-misuse-abuse framework from cognitive science helps design proper human-AI collaboration: proper use leverages human expertise, disuse occurs when systems are poorly understood, misuse comes from over-reliance, and abuse is applying AI everywhere without thought.
  5. The voice AI team created an open-source evaluation framework for voice agents, measuring accuracy and experience, and highlighting challenges like cascade failures in speech-to-text transcription.
  6. Engineers report that AI tools like Cloud Code act as accelerators for code reviews and problem-solving, similar to how calculators enhanced mathematicians' work.

Summary:

This episode explores what it means to be "AI-native," emphasizing that it is not merely adding AI features but building with AI as the foundational operating system. Dai Li distinguishes between responsible AI (governance and accountability), ethical AI (philosophical values), and human-centered AI (designing for human interaction and agency). Dr.

Elena Beaver introduces ServiceNow’s first-of-its-kind open-source AI model accessibility checker, which allows anyone to test AI models for accessibility compliance, promoting industry-wide responsibility. A cognitive science framework from a research leader categorizes human-AI interaction into use, disuse, misuse, and abuse, urging design for proper use to avoid over-reliance or rejection. The voice AI team shares their open-source evaluation framework for voice agents, addressing challenges like transcription errors in cascade systems (speech-to-text, LLM, text-to-speech) versus audio-native models.

Engineers Ian Thurlow and Andrew Yen describe how AI tools like Cloud Code accelerate their daily work, acting as a senior engineer for code reviews and problem-solving, without replacing human expertise. Overall, the episode underscores that being AI-native requires embedding responsibility, rigorous evaluation, and human-centered design from the start, ensuring AI amplifies human cognition and agency while maintaining transparency and accountability.

FAQs

Responsible AI is the organizational and governance layer focused on processes, accountability, and compliance. Ethical AI is the philosophical layer asking 'what should we do' regarding values like fairness and transparency. Human-centered AI considers how humans interact with and are impacted by the technology, preserving human agency.

It is the world's first free, open-source AI model accessibility checker, hosted on GitHub, that allows anyone to test AI models for compliance with accessibility standards and see how they score on a public leaderboard at amac.ai.

The framework covers proper use (effective human-AI collaboration), disuse (shutting off a poorly designed system), misuse (over-relying due to overtrust), and abuse (applying AI everywhere without consideration).

The cascade system uses three separate models: speech-to-text (STT) to convert audio to text, an LLM to process the text, and text-to-speech (TTS) to speak the response. This can cause latency and loss of tone or emotion.

It is a framework that generates simulated conversations (e.g., flight scenarios) and uses metrics to evaluate voice agents on accuracy and experience, aiming to set industry standards for voice agent evaluation.

Engineers use AI tools like Cloud Code to ask questions about code, assist with code reviews, and accelerate problem-solving, treating AI as a powerful accelerator rather than a replacement for human expertise.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.