Go back

How Semantic Layers and Ontologies Create Trusted AI

53m 20s

How Semantic Layers and Ontologies Create Trusted AI

The podcast emphasizes that achieving trusted AI in the current era requires a fundamental mindset shift from traditional rectangular data modeling to network-based thinking, exemplified by ontologies, semantic layers, and knowledge graphs. Jessica Talisman explains that a BI semantic layer focuses on labels and metrics, while an AI semantic layer incorporates controlled vocabularies and definitions to convey meaning to both humans and machines. Tony Seale adds that a context layer goes further by linking these concepts to individual data items through relationships, creating a network that mirrors AI’s own neural architecture. Knowledge graphs, built on open standards, collapse conceptual, logical, and physical layers into one structure, enabling seamless transitions from abstract semantics to precise data points. This approach is critical for disambiguating queries, especially when terms vary across departments (e.g., "revenue" meaning different things to sales and finance). To avoid pitfalls, experts recommend starting with specific, high-value use cases and competency questions derived from business language, rather than attempting to model the entire organization at once. Industries like healthcare and finance have successfully used knowledge graphs for interoperability and discovery. Ultimately, the key to trusted AI lies in disciplined modeling, relationship-focused data structures, and grounding technical efforts in tangible business outcomes.

Transcription

8183 Words, 46093 Characters

English
We are in the age of AI right now, and that's a network shape. If you keep cleaning too tightly to the old fox shaped thinking, then you're unfortunately going to be left behind. Your ontology is like your thumbprint, your digital thumbprint for your organization. It's unique to each organization. And so your mode is your ontology. Hi, data and AI friends. We are at a generational shift. Some might even say a once in a lifetime opportunity. To capitalize on this opportunity, you need more than just great technology. You need a mindset shift. That's why I'm proud to say that ThoughtSpot sponsors this podcast. Go from raw data to trusted insights and action with ThoughtSpot Agentech Analytics, a platform powered by a suite of specialized AI agents. It's why companies like Cisco, Lift, T-Mobile, Sephora, and Schneider Electric count on ThoughtSpot. See what the future of analytics looks like at ThoughtSpot.com. Welcome to the data in AI Teefe. I'm your host, Cindy Howzen. Semantic layers and context layers are your critical ingredients for trusted AI. In fact, some VCs have dubbed them the next trillion dollar opportunity. To help separate the hype from the reality on this critical topic, we are bringing you two world-renowned experts. Jessica Talisman and Tony Seale. Jessica is the CEO and founder of the Ontology Pipeline. She has spent her 25-year career in library science and enterprise AI and building knowledge systems that make machine intelligence trustworthy. Tony is the founder of the Knowledge Graph Guys. He has spent over a decade delivering mission critical knowledge graphs for Tier 1 investment banks. Together, they bring decades of hands-on experience, building the knowledge infrastructure that makes AI actually work. In this episode, we're going to break down the critical components and concepts that you need to have semantic layers, context layers, knowledge graphs, and help you craft a plan for trusted AI without giving away your most valuable IP. Jessica and Tony, welcome to the data in AI chief. Thank you. It's wonderful. Thanks for having us here. Now, we are literally spanning the globe today. So, Jessica, where in the world are you joining us from? Oh, gosh. I'm in Santa Cruz, California. Beautiful. And Tony, what about you? Well, I'm actually down in Devon at the moment. I live in London, but yeah, I'm in from the UK, so yeah. Very nice. And I'm smack dab in the middle on the East Coast in Delaware. So, as your guide today, I want to start with some simple concepts. Now, many of our listeners come from a data background, database, and let's say a BI in analytics world. So, they understand the basics of a BI semantic layer. Jessica, can you start by describing like, what is the difference between a BI semantic layer and an AI semantic layer? It's a really interesting question. When we try to break apart the semantic layer, what is traditionally been defined as a semantic layer, it actually started with business objects back in the 1990s. So, if we look at this historical foundations, it evolved into a metric layer. And where it sits right now is we're starting to see another evolution of the semantic layer, where it includes the metric layer, but we're looking at controlled vocabularies, actual labels, human and machine readable labels, controlled vocabularies, things like that that can tell both AI and humans what the meaning of a word is. So, it simply hangs at the idea of a label. And it can include definitions, but that's where it stops. Okay, thank you. So, and for those listeners that are joining us, perhaps on YouTube video or on Spotify video, I'm going to just share a very basic BI semantic layer. Now, Tony, and let me say Jessica, you use the word vocabularies. So, as we have both Americans and Brits here, so maybe a sample vocabulary where this plays out and Americans might say Ravinil and a Brin might say bookings or turnover. And turnover is a good thing in British English. Did I get that right? Yeah, yeah, absolutely. Yeah, which would be like synonyms. And as well as control vocabulary, I guess there's ontology as well, which then, yeah, you kind of start talking about the more formal structure of how those things would relate to each other. So, you could then get down to here's this concept, but here is his embracation here is in Americans. So, yeah. Excellent. So, thank you for clarifying that. So, Tony, let's peel that onion a little more. Elaborate a little more on what Jessica said. How is a context layer different? What more does it have in the age of AI? Yeah, so, I guess here when we're talking about a context, so we've got the semantics, which is part of the context, but then, you've then also got the individual data items that exist underneath that semantics and how they fit together into like a network of information. So, we could say that the first of all, the AI is coming in and it's needing to understand what a given user is asking. So, a question will come in in natural language and that's going to have to be disambiguated in some way that the machine can understand. And like you were saying, you've got the English and you've got the British, but then each company has maybe got very, it even perhaps even different departments in the company have very specific semantics about what they're talking about. You could be in some detailed aspect of like digital twin, within a factory or you could be well within a marketing department where they've got very specific semantics about what they're talking about here. So, the first part that the AI is going to do is going to go through that semantics part. And it's going to kind of understand what the concept means and ambiguously. But that's not enough because now you need to go from the concept itself down to the data items. So, these concepts, as we talked about, like once we put them into an ontology, they're going to have connections between them. So, like your revenue is related to your profitability or your income, your outcome. These are kind of different concepts and they've got relationships. But then we drill down further into the context. It's like, what is the actual revenue for this particular customer? I've got a customer here, I understand what their concept is, but what's the data for this customer and now what's their precise input and output? And that then kind of gives you this network and that's the context that the AI would need in order to be able to answer your question accurately and precisely. Okay, great. Now, I want because I want to make sure we're really clear here for listeners where this is all new. You also introduced ontologies. This is really how the business works. What is the process from orders to sales, to invoices, issued, to invoices paid? Is that a fair way, a simplistic way, but is that a fair description? Yeah, I think that's absolutely it. I mean, I like to say that you're trying to boil down the essence of what the core of your company is. So, we try to define the ontology in terms of how the business people speak. So, if you were to go into a room of the specialists there within the marketing department, what words are they using? And then those kind of nouns, they become the concepts within the ontology and the verbs that they become like the connections between different concepts in the ontology. And it's the relationship in between. I mean, that's where context really comes from, is the relationships between concepts. And that's the descriptive context. Thank you, Jessica. And you also alluded to it may vary based on roles or departments. So, I often sometimes chuckle and say the definition of sales for a sales person is when the order was issued because that's when they get their commission, whereas sales or revenue to a finance person, particularly if they're in a cash basis, is when the funds are received. So, same, the sometimes same terms, but the context can have wildly different outcomes. So, maybe to share another simplistic concept. layer that we've got metrics, ontologies, context. We're going to pull on the knowledge graph next. We have lots of different sources, lots of different endpoints. And if you're not already following Tony Unlinkton, I really liked this comparison on one of his linked in posts. Now, Jessica, you've written a lot about knowledge graphs and how they relate. So this can look very scary. I'm sharing this image from your sub stack and excellent sub stack. Take us a level further now. What is the knowledge graph? Why does it matter now more than ever? So a knowledge graph represents, on a very simplistic level, the network. Exactly what Tony was describing, the meaning of your data. So it's the concepts and the meaning. But in fact, a knowledge graph is like an architecture, much like we have a data architecture. So what we're really doing when we create a knowledge graph is we're structuring the concepts, the relationships between the concepts, and we're modeling the data. So it's a layered approach that semantically structures things in what we call using ontologies, what we call triple statements, which are subject object predicate statements. And that is the descriptive layer. So that's really, I think, what is new to a lot of people is that where we're not just hanging out with labels, but we're reconciling in order to disambiguate, in order to describe the things within our institutions or our environments. Yeah. Well, is a former English major subject predicate? I'm got it. Yeah, got it. Now, because I have to make sure I'm bringing our listeners along here as I'm going to play their role, one person said to me, why are we making all these fancy new terms? If it isn't this really logical data modeling, agree or disagree, Jessica? I would disagree, because in fact, within knowledge graph work, we have the conceptual layer, the logical layer and the physical layer. And so it's a different way of engineering. I think what is most important is to realize that these are open standards-- based on open standards. It's a specific way of modeling. And so it's known as symbolic AI. We'll leave those knowledge graphs. So the conceptual layer, in fact, is our vocabulary. It is what we know as sort of the semantic layer within our data environments. And then-- so that's our conceptual layer. The ontology we can think of as the logical layer. So that's correct. OK. But then, a knowledge graph exists as the combination of both. Got it. And it's not new. It's newly hyped, maybe. Newly hyped. Newly hyped. Which industry is most advanced in using them, Jessica? I think health care. And I would say second to that is finance. But health care absolutely has relied on ontologies. In fact, a lot of the genomic work that has been happening successful genomic work with AI has relied on ontologies. And the community has really rallied around in order to create these layers of not only transparency, but interoperability. And I know that's a big keyword with semantic layers. It's part of the allure of semantic layers. But we see that come to fruition with knowledge graphs. It's really built first and foremost to support interoperability. And so when we look at-- when we look at sciences, that interoperability is critical for research and discovery and innovation. Yeah, research, discovery, the provider to pay our continuum exactly. So Tony, any threads you want to pull on there, any clarifications in the concepts that you want to elaborate on? Yeah, I would. Two important things to pull out with the knowledge graph, with the kind of pushback on the logical, the model thing, which I think is interesting. And we do-- and up across that, when we're doing the transformation programs with customers. So I think there's a couple of things to understand here. So one of them that Jessica brought out before-- an knowledge graph is different because it's a network. So we're not talking about sets here of putting stuff into these rectangular box shapes. That's what we've done all of these years. I need to model something now. I'm going to get a rectangular tabular shape box, and I'm just stuff, everything into there. And when you're thinking in that mindset, it's always the data items. It's always the entities that you're concentrating on. And you don't really think about the relationships between those entities. Or you have some joins between your tables, but they're not first class citizens. And part of shifting towards a knowledge graph is that the relationships between things becomes as important as the things themselves. So we're thinking about our data in this network now. So your friendship network on Facebook or something is like a class example. But you can do that to any data. And AI loved that, basically, because AI is itself a neural network. So it's in that same-- you put your data into the same shape that AI is already in, which is like a network. That's one thing to say. The other important thing to say is, again, with traditional data models, we've tended to split the schema from the data. So you've got the schema, which is in its own strange format often in these additional tables. And then you've got the tables themselves. Within a knowledge graph, you pass from conceptual to data. It's all just the same structure. It's all just part of the network. So I'm able to go up into the abstract conceptual layer and do abstract thinking, is it well, with the aid of a large language? And then I'm able to come down into the data points. And I can just transition between those two things because they're all in one networked structure. And how do this manifest yourself? Well, with clients, you sort of go, OK, logical, to call conceptual, physical model. And you kind of have to end up sort of somewhat collapsing those things when people are talking about. We want to start with that conceptual. So with that conceptual for us, that's the semantics now. And we want to be making sure that the semantics are as close to the reality of how business people are speaking. And then once we've got that, we call that an ontology. And then once we've got the ontology, we connect it into the data. But it's still within the same network. And that we're kind of physical, more or almost kind of disappears at that point. It is a transition in thinking. It is a new way of things for those reasons. Got it. Thank you. And now, because my brain is picturing these connections, so organizations, there's different levels of a gentick. And some are trying to get to agent to agent communication, just within their company. When we go cross company, that's even more complicated, bolder, few are there yet. And so if we think of these as agents as building blocks to the autonomous enterprise, would it be safe to say, you've got to start with just the relationships, the graph, even for just a small work process, before you go this whole, how does my entire company work? What do either of you think? Well, I think it's a discipline. The idea and parlaying off of what Tony just said, when we see all of these layers collapse, it's learning how to model first and be able to define data and connect into relationships instead of jumping to the end state. I like to say if you can't enable a simple chatbot, then you haven't really arrived there yet. And so yeah, yeah. OK, but Jessica, you're saying model first. I feel like a lot of the people that I work with are like, just give me a chunk of data. And let we have at it. Am I working with children here or what? A little bit. I mean, because in order to define and to reconcile semantics throughout an organization, there has to be strategy. And part of that strategy is determining what your approach is going to be. And that's where it's difficult to hold the horses from the barn because we are so anxious to arrive at this end state without doing the hard work of defining and modeling our data in a way that's meaningful that we can actually interoperate within our own organizations before turning that outwards to other organizations. Yeah, fair. And there is so not everyone is an ill-behaved child. There are so excellent data modelers. And this is where experts like Joe Rice and Sunny Rivera and Mike Ferguson and we would debate like why did data modeling go on decline? And I think I would blame some of the visual based data discovery tools where you didn't need it. You just had a dump and could go from there. So I think in AI and a chatbot world, we do need that discipline. So some people are going to either dust off their skills or get some new modeling religion. So let's make this real. Tony, if you can maybe take. in industry use case, whether it's pharma or as you work with some tier one banks, what can you share about what good looks like or what missteps to avoid? The first one were probably related to what you were talking about there. And Dave McLean has got a good phrase which is thinking big but delivering small. So I think there's a lot of work to do here but starting straight from, are we going to model out the autonomy for the entire organisation? That's an overdaunting task that's not really going to work very well. So what you need to do is you need to ground that ontology into examples. So you need to select a use case that's going to have high value, that's going to have good ROI and you can prove that it's going to do that. Because the classes in the ontology, they don't actually have any value in themselves. There's no point just kind of building an ontology for the sake of it or knowledge graph for the sake of it. You're doing it for business reasons. It's like the kind of classical stuff you want to identify those use cases first and be sure that they're providing value. Then what we tend to do from there is we generate these things called competency questions. So like, okay, what questions would you ask in order to be able to resolve this use case? And from that, if there's competency questions are written in the semantics of the business language, then derives the ontology. Because the ontology could be anything and it's quite possible to get lost in this kind of philosophical world. Whereas really what you want to be doing is concentrating on modeling out the semantics, which is going to affect your bottom line of your company. So classes related in your ontology should be related back to questions that you can answer that have high commercial value for you. So I think that's one thing. And I point people towards the D-project specification, which is the semantic data product specification, that a number of us got together and didn't came out of the financial industry, which is an attempt to kind of do that, to tie data products to classes within our ontology so that you've got that direct linkage there. Between a direct product which is service me use case and the class in the ontology that it's supplying. That's one. Another one, which classic thing is when people think about noise graphs, then often the first thing to do is create a graph visualization. It's just a very natural thing. But the reality is that's often not the most useful thing about a knowledge graph is having a graph visualization. They can sometimes be useful, but quite often they do this thing of creating what a friend of mine called the hairball, the tail. Basically, it's just, you've got so many, it turns into this amorphous blob, and it's not actually very useful for decision making. So yeah, it's important to realize that knowledge graph is like way more than just a visualization of that graph. It's more like the kind of electrical wiring in your house, basically. You don't need to see all of that electrical wiring in order that for use for you just need to press the switch and the light turns on in your room. Yeah, you've got to complicated electrical network within your house, but that's sort of a single use case. So in some ways these two points interrelate to each other. I'll just say one more, I can carry on down for ages, but I'll just do one more. Connections in the data, they follow connections between people, and we're all kind of used to working in these silos and not communicating well with each other. And in large businesses, it's often very fragmented, and people need to come together in order to make this happen, and that's like an organizational transformation process. That's as big as the technical challenge. When you build for purpose, because you're building a model that is meant or intended to answer business questions, then you're modeling, that's where data modeling becomes ever so much more important. Because you're designing and architecting a system with an intentional output, an intentional end state in supporting the business. Which may or may not be too narrow. So when you talk about hairball hell, it sounds like transformation hell too, because even if you take an element like customer, well, a customer could be a prospect, but money is not yet changed hands. Or they're in a community, but they're not an active customer or inactive customer. And this is when you hear the silos say, "Oh, well, I didn't know you were using that data element in that way." So Jessica, you have a case study that you're allowed to share. Can you tell us about this case study? Yeah, broadly, I left Adobe in last summer. And I was a senior information architect there and helping to architect for the entire DX side of the business, which is like Adobe Analytics and AP and sort of the business products, not the creative products. But it was a very interesting problem. And this parlay is off of the customer example. And most organizations struggle. It's not too uncommon. The connection between pre-sale and post-sale. This huge gap in understanding and language. And it ultimately is meant to support customers through their entire journey. So an example is just what we term the semantic layers, reconciling the vocabulary being the first order of operations is that we have to all agree upon terms. Or we have to leverage ontologies in order to connect disparate vocabularies so that we can start to go on that journey. So Tutonis' point, when he was discussing how the language, the silos, how this is so much a sociotechnical problem, just starting with the vocabulary in the semantic layers, a great start, but it's not the end goal. So to really connect the meaning and to be able to pick up on the red from pre-sale, someone has then engaged a company has decided to purchase an Adobe product. We have to be able then to follow that customer journey all the way through product usage and the implementation of their instance, how they're using the product and what type of support they need ultimately. And obviously they're not enough people in the room in order to support a lot of these customer journeys. So we leveraged ontologies in order to support these customer journeys so that we could continue to support customers after they signed the deal, assigning, for example, solution consultants that were part of the post-sale movement in order to meet their specific needs with the product. And that included, and this wouldn't maybe seem as logical, but some things emerged from that activity and that work in which there was a realization of needing more learning support, educational support for some of our post-sale customers. And making sure that that message was received by the post-sale solutions consultants in a timely manner in order to maintain communication with customers. We also supported with the entire learning environment then using our knowledge graph that we built in order to tag content appropriately, so building an auto-tagger, so documentation and learning objects could be identified and assigned appropriately to customers in the post-sale motions. So that is sort of, you know, before I think the before state is stitching. We just try to stitch things together within sort of SQLized environments. And, you know, now with AI, I've been able to leverage AI as a partner, an appropriate partner, in order to do so accurately. It really, you know, that that shift started happening where we started implementing our ontologies in our knowledge graphs. >> Yeah, thank you for sharing that. So you use the word stitching together. And is this a little bit the state of the technology? And or the state of the mindsets and skills. So if we get really specific, there's a new concept in our industry, the open semantic interchange led by Snowflake with many founding contributors of which Thought Spot is one, you also have some of the observability of your catalog vendors, DBT, Atlin, supporting it, Databricks recently joined it as well. But what should people be aiming here? What should people be aiming for here? Is it one artifact for each of these components? Or is it a stitching together and interoperability? Is stitching together is, I think, old hat. It's something that we're used to doing. It's a muscle that many of us are familiar with. And it happens. I see it most often as complex SQL queries and joins. And sometimes you can hand mapping between objects within data tables and what have you. And that in and of itself is not semantic. And that may seem salacious. But then there's the other issue of is YAML semantic to the OSI point. And this is another interesting experience that I'm going to talk about. Adobe that I personally dealt with was having to reconcile, having YAML config files. And what's really interesting is when we take something in a knowledge graph sheet or you know take ontologies and then we try to flatten it into YAML config files in order to push it into production, it actually stripped a lot of the semantics out what we delivered with ontologies. So I think that really it's about upskilling. It's about, you know, we've been in this sort of data only, you know, dealing with relational databases where those, you know, relational databases have a originally designed for storage and query. And so the bigger question is is in order to imbue semantics within a system, it's moving beyond that syntactic view and being able to combine approaches, hybrid approaches. So that does require skill building to break people out of the old ways of doing things and start to embrace the more descriptive aspects of taxonomies ontologies, those sorts of things. Yeah, thank you. And Tony, that ties into you wrote the term in one of your posts that we have to overcome deeply ingrained relational database thinking. So it sounds like you and Jessica are aligned here. Yeah, I mean, let me preface this by saying that it's important to be pragmatic here. So relational databases are not going anywhere and nor do we want them to go anywhere. They're a good piece of infrastructure. And when we're putting these semantic layers in, we're, they can, it's called a layer for a reason because we're layering it as an interface over the top of existing infrastructure. You get to keep existing infrastructure as it is and you'll probably carry on like that for an awful, awful long time, I imagine. But just like Jessica saying, there is a compromise to be made and OSI is a compromise. I mean, we use OSI for those pragmatic reasons. We use it in our implementations. And I think it's a good positive move actually that if for the industry and I fully support it. But it's not true semantics still. It's because of the reasons if we refer back to what we were talking about in terms of needing both the network structure for the kind of rich connectivity. And also even on the semantics level, like if you were kind of taking your data and you're putting into these, these kind of like rectangular shapes, you're, you're compressing the semantics. So what do I mean? But well, in an ontology or even a control vocabulary, you have these like rich hierarchical structures. So you know, you talked before about like, oh well, a customer, a sale might mean different things in different departments. Well, that would mean that you could have a quite an abstract concept of sale that share between those departments and departments on what they can agree on with the properties in there and the descriptions of the semantics. The what a sale is in general. And then the different departments could have subclasses so they could inherit from the kind of more abstract version of sale and then they could. So if you think about like the animal kingdom, for instance, the way that we kind of have like sort of animals, their reptiles and mammals, and this is like what we call like a taxonomy in that sense. So that exists within the rich semantics of other ontology, all of these things and not even just at the kind of class level, but because relationships first class systems, you can have these inheritance hierarchies within the relationships as well. So that gives you an awful lot of flexibility for modeling there. When you come to put that into a table, what have you got to do? Well, I've got my whole animal kingdom and I've got one animal table now. So in a star schema, I've got to flatten all of that rich complexity of that rich semantics. It's good. It's useful. It's not going away anywhere, but let's be very clear. It's the hallmark of the industrial age that we are transitioning out of now. Like in the industrial age, it was like the machine metaphor at that point and we built these boxes in everything. Just go kind of look at the other world that we created was all industrialized farming with these kind of like very straight linear lines and the buildings that we created, this kind of like that was the industrial age, but we are absolutely passed through that now into a different period of time. We're in the age of AI right now and that's a network shape. So if you keep clinging too tightly to the old box, wet box shaped thinking, then you're unfortunately going to be left behind. Yeah, that's a fair point, but I also like your point about being pragmatic because we don't get to throw away both the technology. We need to evolve the skills. I don't want to throw away too many skills, but we need to evolve the skills. But I think the network aspect is real key. I want to shift to a slightly different topic and this is what happens when we strip out the origins of this information or or provenance as some people would call it. Jessica, what happens when we strip out these citations when we're using AI? Yes, that's that's very important because there's two layers essentially to a model in LLM and we have the parametric layer and that's where the training data, your model arrives and out of the box. It's already been trained on the internet. And for that matter, most provenance has been stripped out of the actual knowledge that is held within these models. And so we have the retrieval layer and that's really what we're all grappling. So with the semantic layer, these are how can we add context? How can we add context? And so we have things like RAG and other types of implementations. Now that is our one chance in order to inject or to describe the content or to model context and provenance. Knowledge and our very knowledge ecosystems, how knowledge has evolved over thousands of years has been reliant upon provenance. So it's the lineage of things. It's the idea of how we build on top of the shoulders of giants, how research happens, how innovation evolves is reliant on provenance. And what's happening at a very large scale is LLM's that can plagiarize and not necessarily tell you that they are plagiarizing. And so there's an issue where we receive knowledge from an LLM or we are modeling knowledge and forget to include provenance. And we strip apart that very important lineage that tells us how things or ideas or knowledge has evolved over time. And so yeah. Well, so I mean this opens a whole so many things. Do I trust the source, but also do I credit the source? Who did this pioneering work? So there's a lot more to it than just was it a good answer? And it's in library science, which library science has been instrumental in modeling knowledge and maintaining these repositories of knowledge for centuries. And the entire ecosystem, which also is available in sort of ontologized, those are knowledge ecosystems and library ecosystems. So catalogs have ontology backbones and connect to very familiar sources like Wikipedia and Wikidata. Google Amazon. So a lot of what we're trying to leverage, it's known in library sciences, scholarly communication. And it's an actual domain of study and it's a domain that there's centuries of work around this is that it's essential that we break down these silos, but it's essential also that we maintain that lineage because it's only the accuracy of the information you receive is reliant upon the lineage or the provenance of the information received from an LLM. It's so important. Yeah, but let's let's also say I'm a business leader and I'm busy and I just want an answer to my question. So how should leaders think about which LLMs or what set of capabilities they should be looking for? Yeah, so that really starts from inside because when you receive a model, it's just going to come pre-packaged and it has this and lets you build your own smaller medium-sized model and you have complete control. So it's really about implementation and how you choose to to architect your information ecosystem or your knowledge ecosystem. So it's including things like provenance, even if you're modeling the semantic layer, you have to have those signals available and provenance is reliant upon the relationships between things. So there's no way to model and to encode provenance without including the relationships between business objects even. And it's important for the leader and that's the other thing is that I would hope that leadership would want accurate information and not to flood their networks with misinformation and disinformation. So it's really is in the best interest of a business to also invest in that and to invest in knowledge management and we haven't really gotten there yet. Yeah, I would say they, they're not going to use AI unless they can trust it. And it's how do you build that trust when so much is being reinvented? Let's shift to another important topic. And this is the anthology as a company's core IP. And Tony, you have written about this. You've warned people about not giving away their IP. Tell us your thoughts here. Yeah, so I guess if you, if you follow the thesis, which is the basically during this phase of the AI transition, I like to talk about the AI iceberg. So it's don't focus on the very pinnacle, which is the top part that that's the algorithm, that's the model. That's not what you need to be concentrating on during this period. As a business leader, you need to be looking below the surface to the data infrastructure. So the key trick to do right now is to turn the power of the models that we've got back upon your own internal infrastructure in order to build out these rich ontologies and to connect your information together. That's that's if you like that the move to do. And when you're doing that, what you're going to be doing is you are going to be looking at the information, which is out of distribution with the wider information, because as Jessica says, the foundational models, they basically took all of the internet. So we the people put our information out on the internet and connected everything together with with hyperlinks. We created a big graph, basically, of everything that we know. The hyperscalers, the foundational high scalers of of taking all that information, they've compressed it down and built this, this incredibly powerful models, which can now do all of this stuff. But that's the kind of distelt information that everybody already knows. And then anything that is in that space, the value of it has just dropped to like 0.1 cent or whatever it costs you in order to make a large language model call. And whatever you do that's in that kind of wider distribution of information, the large language models are now going to be able to give you a very accurate answer to. So they're highly unlikely to hallucinate if something is like well known in the public sphere because they've seen so much training data about it and they've been fine tuned on it. So they're going to give you an accurate answer. But the second that you step out of that space into something that's out of that distribution, that's when the large language model kind of it goes from being a genius to being like sub five year old and just one move it falls off a cliff basically. But there's another thing that happens in there, which is the value of that piece of information has just shot up because it's no longer in that I can just get this from a large language model. So if we return that back now to a company and the moves that I believe that all companies need to make at this point, you're taking the power of the large language model and you're turning that back internally to reflect upon your organization and you're saying what data have we got that nobody else kind of knows about now and what are what what tribal knowledge is inside people's heads about what's important of how this interoperates. The tribal knowledge we want to encode that down into formal semantics into first order predicative logic within a proper ontology and then we need to do the other thing which is we need to connect the information together. So all these separate systems we need to connect them together. The causes we talked about before ontologies have got this ability to do abstraction so to create more and more abstract versions of the concepts, different departments that are doing different things. Eventually what you're trying to do is you're trying to consolidate out the core abstract concepts that are fundamental to your business. That then as you can kind of like distill it down and learn it over a period of time with a hell of a lot of hard work. There's no sugarcating this pill. It's a huge amount of work in front of pretty much every medium-sized organization and above globally around the world right now. That's what we talk about. It's a trillion dollar market in all of this business so wake up and use there's a huge amount of work to do and AI is not going to take it all away. These are tools that we can use but you need humans in the loop here. What that will do eventually hopefully is that we'll consolidate and compress the unique ontology for your company. That's what puts you out of distribution. That's where all of your value lies. Do not repeat. Do not just give that away to somebody else. If you go left somebody else, take your ontology, host it on their system, learn the essence of what is that you know that's out of distribution with the rest of the world. You've just given them a head of think this value of your company on that. That's a really good warning, Tony. Jessica, you also have warned people not to rely only on the public ontologies. Do you want to add anything there? Yeah. I mean, I like to say that your ontology is like your thumbprint, your digital thumbprint for your organization. It's unique to each organization and what how you define things may not be the same as an Ellen might define something. Going back to my Adobe example is there's feature in almost every single business product called journey. And so defining that uniquely as a product feature and not actually a journey or is it a customer journey? I mean, you can imagine all the different ways that we can define and add meaning to the idea of a journey, how your business models that it's instrumental that that and I agree with Tony that you guard it as your own IP and it is your differentiator. So we always talk about what's what's the moat? I hear that especially a lot for the B-seeds and so your moat is your ontology. Got it. Thank you. Yeah. You know, so you both have described some complex topics in just under an hour. That's all we have here. And the criticality of getting this right for the AI age. So I was excited to see you recently launched the knowledge graph academy. When did you decide to launch this and what is it? We've been feeling in the in the semantic and ontology space. We've been feeling generally as a community that there's this huge gap obviously between relational database and traditional data management, methodologies and you know, ontology and knowledge graph community. There's a need to build bridges and a lot of those bridges, you know, it can't just simply be implementation or trying to sell a product, you know, the vendor's space. It really involves skills and I think that's what's emerged from our conversation today and what we felt like was sorely missing. So it was in the fall and we sort of got together with Katarina Kari in Finland and started this initiative to start the academy as a way to provide like meaningful education for people to so that they can pursue these skills and actually implement knowledge graphs in their own environments. Amazing. Tony, anything to add? Yeah, not just that I'm incredibly proud of what Jessica and Kessarina doing with the knowledge graph academy. It's an important effort in order to be able to uplift people's understanding like ontology, like all the stuff that we've talked about just now, they're complicated words, they have complicated meanings. So yeah, to have a program out there where seasoned experience experts are able to kind of pass that knowledge on to people. I think it's a really important part of making this kind of transition successful for a lot of people. We actually want, I believe, as many companies as possible to adopt this and be successful in this in order to so that we still keep a diverse of companies on the other side of the AI transformation as we go through it. The other thing is that we've graduated now 23 students and we currently have another 20 odd students in flight through our program and we're starting to stand up in-house so the customized trainings within organizations. So back to the idea of change management is how do you get towards change management? It's really education. We're not guarding these skills. The idea is to democratize these skills so that we can empower people to move through this transition rather than feeling disempowered. Excellent. And for me, as a podcast host, what's really fun? I said, we need to do a podcast on this topic. I'm like, I like these two people. They're work. I like how they explain things and only afterwards. Did we find out you were already working together? So that is just magic. We're going to shift now to a fun lightning round. Short rapid fire questions and answers. Jessica, when you're not working in context layers and knowledge graphs, what do you do for fun outside of work? Oh gosh. I love gardening. So I have quite the garden and I spent quite a lot of time and I think it's perfect because Santa Cruz. I live a few blocks from the beach, but Santa Cruz is such a temporary climate. So the gross season is all year round. So that and yoga. I've been practicing yoga since my late teens and I'm actually also a yoga teacher and so that's kind of an outlet away from tech for me and where I find my center. Fantastic. Tony, somebody who inspires you famous or not famous? I'm going to go for Tim Bernadilly, actually. He's a hero of mine. So he invented the web, but then also his kind of like second thing that he was part of venting was lightning data and the semantic web. which is basically the technology that we are advocating for here. And to do it in such a way that it's in open standards and it's for the public, in a public commons, I really admire it. - Okay, love it, thank you. And then the last one, Jessica, you get to choose this one first. Either what are you most grateful for, maybe be on the obvious of health and family, or something you've accomplished in the last week that you are particularly proud of? - I think I am most grateful for the ability and it says beyond health, is the ability to have these exchanges. I am so grateful for community and the opportunity to be able to discuss these things to help too. I don't think if I were to look at myself a year ago, I didn't have the same networked community because everything was so rife with, I think, anxiety for everyone of where the industry is going. But I think that we've really sort of developed this interesting community and ability to exchange at least past domains. So the ability to engage with those conversations and be a part of them, I'm very grateful for that. - Yeah, as am I, thank you, Tony. - I'm grateful to be playing a part in this point in history. It seems like one of the most exciting times that we're going to get to live through and there's so much to play for and things could go really well and we could get really good AI or things could go really badly. So having the opportunity to stand up there, I guess I'm grateful to myself that 10 years ago decided I was going to commit to this aspect of knowledge graphs and AI as it was back then to commit to doing that regardless of whether it was popular or not, but just because it felt like the right thing to do and all of that experience that I've now been able to accumulate in that period feels like I've has prepared me for this moment that I can now stand up and play my part in one faults. I'm just super grateful to be in the position that I am to do that. For sure, you both have done pioneering work a little bit ahead of its time. And so we're grateful that you're sharing your insights, your expertise. I'm also super proud of the accomplishment that we got both of you at the same time. So wonderful to have you on the podcast. If you've enjoyed this conversation, please rate or like it on your favorite podcast platform. And if you would like more inspiration and thought leadership, follow me on LinkedIn or visit thedataandaicheast.com. [BLANK_AUDIO]

Podcast Summary

Key Points:

  1. AI requires a shift from traditional data modeling to a network-based approach, where relationships between data entities are as important as the entities themselves.
  2. Ontologies serve as unique digital thumbprints for organizations, defining controlled vocabularies and concepts that are both human- and machine-readable.
  3. A context layer extends a semantic layer by connecting concepts (via ontologies) to actual data items, enabling AI to disambiguate queries and provide precise answers.
  4. Knowledge graphs integrate conceptual, logical, and physical layers into a single network structure, allowing seamless transition from abstract semantics to specific data points.
  5. AI thrives on network-shaped data, making knowledge graphs and ontologies critical for trusted AI, as they align with the neural network architecture of AI systems.
  6. Successful implementation requires starting with high-value use cases and competency questions, rather than attempting to model the entire organization upfront.
  7. Industries like healthcare and finance are leading in knowledge graph adoption, leveraging them for interoperability, transparency, and innovation.

Summary:

The podcast emphasizes that achieving trusted AI in the current era requires a fundamental mindset shift from traditional rectangular data modeling to network-based thinking, exemplified by ontologies, semantic layers, and knowledge graphs. Jessica Talisman explains that a BI semantic layer focuses on labels and metrics, while an AI semantic layer incorporates controlled vocabularies and definitions to convey meaning to both humans and machines. Tony Seale adds that a context layer goes further by linking these concepts to individual data items through relationships, creating a network that mirrors AI’s own neural architecture.

Knowledge graphs, built on open standards, collapse conceptual, logical, and physical layers into one structure, enabling seamless transitions from abstract semantics to precise data points. , "revenue" meaning different things to sales and finance). To avoid pitfalls, experts recommend starting with specific, high-value use cases and competency questions derived from business language, rather than attempting to model the entire organization at once.

Industries like healthcare and finance have successfully used knowledge graphs for interoperability and discovery. Ultimately, the key to trusted AI lies in disciplined modeling, relationship-focused data structures, and grounding technical efforts in tangible business outcomes.

FAQs

A BI semantic layer started with metric layers from the 1990s and focuses on labels and definitions. An AI semantic layer goes further by including controlled vocabularies and human- and machine-readable labels to convey meaning to both AI and humans.

A context layer adds individual data items and their relationships within a network, going beyond semantics. It disambiguates user questions and connects concepts to specific data, like revenue for a particular customer, for precise AI answers.

An ontology is a logical layer that defines the concepts and relationships within an organization, using business language. It acts as a unique digital thumbprint, representing how the business works, such as the process from orders to paid invoices.

A knowledge graph is a network architecture that structures concepts, relationships, and data using ontologies, making relationships as important as the data itself. It matters because AI operates as a neural network, so knowledge graphs align data in a network shape, enabling better understanding and interoperability.

Healthcare is most advanced, relying on ontologies for genomic work and interoperability, followed by finance. These industries use knowledge graphs for transparency and research discovery.

Avoid trying to model the entire organization at once or creating graph visualizations that become a 'hairball.' Instead, start with a high-value use case, generate competency questions from business language, and ground the ontology in specific examples to ensure commercial value.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.