Go back

Data Privacy in an AI Landscape

0m 0s

Data Privacy in an AI Landscape

This podcast episode features Sean Falconer from Skyflow discussing the growing challenges of data privacy and AI. He explains that despite massive investments, data breaches are rising because companies often treat data protection as a secondary concern, not a core business function. This leads to sensitive customer information being scattered and replicated across thousands of systems, creating an unmanageable attack surface that traditional security tools cannot adequately protect. Skyflow's solution is a data privacy vault—a centralized, secure repository for sensitive data. Instead of storing actual personal information in multiple places, companies store anonymized references or pointers to the single, protected copy within the vault. This simplifies governance, reduces risk, and is based on a model pioneered by large tech firms. Skyflow initially markets to industries with strict regulatory requirements, such as finance and healthcare, where the need for robust data privacy solutions is already well-understood and critical for compliance.

Transcription

10392 Words, 56671 Characters

English
Welcome to Ahead of the Game, a podcast brought to you by the Digital Marketing Institute. I'm your host, Will Francis, and in this episode we talk to Sean Falconer all about the data privacy challenges of using AI and keeping customer information safe. Sean is head of marketing at Skyflow, the California-based tech company that provides a data privacy vault service that we'll hear more about when we talk about customer data, privacy and regulation. Sean's from a tech background. He's built and sold a hiring platform startup before spending time working at Google. He's also a host of the Software Huddled podcast. Sean, welcome to the podcast. Let's start by hearing what exactly Skyflow does. I think the easiest way to understand Skyflow is kind of the take a step back and look at the actual problem that we're solving. So over the last 20-plus years, any interaction that you have as a consumer with a business, nearly any interaction, whether that's like going and, you know, picking up pharmaceuticals or booking an airline ticket or buying something online, you're giving up essentially personal information to those companies. And we sort of blindly trust that those companies are doing a good job of actually protecting the information. But if you look at the data, there's been 2.6 billion personal records leaked in the last two years. And this has actually increased 20% in 2023 versus 2022. So it's clear that businesses, despite spending billions of dollars on trying to solve this problem and lock down these systems, are not doing a very good job of it. And I think that naturally makes you ask the question of, you know, why is it that these companies struggle so much with these different problems, even though they're really well-resourced, have lots of technical talent to potentially address this issue. And I think there's a couple of different reasons, but I think the simplest reason is, for most of these companies, it's simply not their job to essentially focus on protecting the privacy of their customers. You know, if you work for Uber and you're hired as an engineer, what your focus is is essentially how do I, you know, deliver a driver to pick up a rider as efficiently as possible or something like that. And that's what you're hiring your tech talent for, and you're essentially, most companies are hiring their talent and also putting those resources behind things that are customer-facing that deliver our ROI for the business. So historically, it's been hard to justify our ROI from like a data privacy standpoint. And it's kind of similar to, for you, you probably don't take the money that you earn in your job and stuff it under your mattress. You know, because if you did that, you would take on the responsibility of protecting that your money. And you probably don't have the resources necessarily the experience to do that. So what you do, typically, is you probably trust a bank to actually protect it, which does have the resources and expertise to do it. And where is your money going? It goes essentially into a vault. So essentially, the idea behind Skyflow is that we provide the technical sort of equivalent of a vault, what we call a data privacy vault, to customers that especially design technology to actually protect customer PII. So rather than essentially taking on responsibility of business, of having to build this, you know, infrastructure or figure out a way I can hire to actually protect the information, I can offload a lot of those responsibilities to a company that's solely focused on essentially solving this problem. And we do that essentially through this concept of a data privacy vault, which is also the recommended approach by the IEEE and one of the approaches that is used widely by some of the top technology companies in the world. But essentially, we made that available to every company in the world. PII, personally identifiable information? Yeah, so another way to think of this is essentially sensitive customer information. If you're a business, what are the things that you wouldn't want showing up on the front page of a newspaper because of the data breach? So there's regulated information that, you know, is essentially regulated under certain privacy regulations around the world like GDPR and in California, there's CCPA, in India, now there's DPD, there's all these different laws and regulations. But even more broadly than that, there's certain information that's not necessarily regulated that you necessarily don't necessarily want leaking up there to anybody on the internet about your customers. It makes it PII, like is the fact that I'm male PII, or is it something deeper than more specific than that that makes it PII? It's really about how clearly you can identify the individual. So male versus female, like if I just knew that you're a male, I probably couldn't purely identify that this is will, but I could maybe combine the fact that you're male and you live in a certain location and that you have glasses or something like that. You know, I could essentially cobble together multiple pieces of information and potentially identify someone and there's actually something that happened several years ago where Netflix had put out, this is kind of like a famous study. They take in a bunch of anonymized data about reviews of movies and they put it up there and they put out this test of they wanted people to figure out like how can we essentially create like a better recommendation system and it was just open challenge. But people were actually able to take those ratings and combine it with things from IMDB where people were actually not anonymized, but rating things and they were able to create correlations where they actually able to re-identify a whole bunch of different individuals. So things can get really complicated, but for the most part, most of when it comes to like things like data breaches and regulations is not as complicated as that is really down to like someone's address, their phone number in the United States, like their social security number, maybe their credit card number. Things that are like clearly identify somebody, you clearly do not want just showing up on the dark web somewhere for anybody to have access to. It's funny, isn't it because you think the insensives already there? I mean, like I don't know if you've ever looked at the GDPR fine tracker website, you can go and look at any of the finds doled out by the EU and see how big the fine was and details of the ruling. And I mean, some of them are billions and billions of euro. I mean, you know, you think that was enough of a disincentive to not have your data breached and that they would put iron walls tripled and quadrupled times around that data anyway. So I think there's two problems there. One is that GDPR, despite like I think a lot of people now being like understanding what that is or at least having some familiarity with the acronym and that it's a privacy regulation to be aware of, it didn't come into effect until 2018, which is not that long ago. So, and a lot of these finds are fairly recent things and they've also targeted a lot of like big companies like Meta and so forth in Google. So if you're a smaller company, you might feel like, you know, I'm not at that level, maybe they're not going to come after me. But then also it's just a fairly recent regulation. It's only something that I think now actually is becoming something that's a high priority for companies, but this is all fairly net new. And now there's over a hundred privacy regulations in the world. So this is becoming more and more of I think like a board from room sea level conversation. But the other problem that companies have is even if they've reached a point where they're like, oh crap, you know, we don't want to show up on this. You know, GPR finds you website and and do damage service is not like any of these companies are like, you know, twiddling their thumbs and being like, haha, we, you know, had another date of reach this week. It's really bad for their business from a trust standpoint and from the fines. But it's still a hard problem for most of them to solve because they essentially I would say taken applied like historically with applied sort of 1980s level thinking about data. Where we've really treated all information the same. So if you go back to like the 1980s when we were first bringing desktop competing into the workplace and people were sort of transitioning from like a paper world to someone of a digital world. At that point, it didn't really matter that I treated someone's, you know, personal information as well as my application data is all just, you know, one zero is like put them in, you know, wherever I'm going to put them because it's basically housed within this can self contained machine where someone actually have that physical access to like get access to an information. So the scale of the problem was much, much smaller, but we've taken that sort of concept of data where everything's just holistically the same we're all going to kind of just put it in a box somewhere a database in the sort of the world of the cloud, maybe put some security fencing around it and so forth. And we've done that at the scale of the cloud to millions and billions of users. And now what is that actually happened is most of these companies are not looking at just like how do we protect essentially the data of where it's house. It's housed in thousands of different locations. So people end up creating hundreds of not thousands of replications of all this information in the database, log files, backups of these systems, the warehouse. It's essentially strewn all over the place. It's a little bit like if you took your passport and you made 10,000 copies of your passport and then you put them all over the place and then trying to protect those locations. That's, you know, a much harder problem to solve. And that's basically what is facing most companies today. And that becomes a really intractable problem because even if they have really good intentions, they end up trying to like switch together through like duct tape and chicken wire, a bunch of different security tools to patchwork this thing together. And you're just never ever going to plug all the holes and you end up increasing essentially the complexity of the system and the technical debt that you're taking on. So you have this like kind of huge like unmanageable system of where data is all over the place. It's hard to know what you're storing where it's stored. And then you have this like patchwork of different security tools and maybe some in a house homegrown stuff that you have to like manage and it just becomes like a nightmare to try to control and inevitably you have some sort of back door or a place that you miss a log file somewhere or someone gets access and then, you know, you're end up in the news. What is the most common way that people get hold of data illegally? What is the most common kind of breach? What's the most common whole, if you like? A lot of the times there are a lot of attacks are like opportunistic, although recently we've seen some pretty sophisticated attacks on places like MDM and Cesar retool where people are combining things like known exploits but social engineering and so forth. But a lot of times social engineering being where I basically, you know, I call up a customer support person I can like, you know, through some form of manipulation convinced them that maybe I work for that company as well. And they I need to like reset my password and they, you know, they essentially send something to me that I shouldn't have access to which thing gives me access to your account and then from there I can maybe leapfrog to access to other people's accounts and so forth. Now, of course, like in the world of AI where we can fake someone's voice and even some of the deep fake stuff around like video, you know, is this really Sean or is this somebody else? Yeah, it's getting really potentially sophisticated and a little bit scary, but usually they're not that level of sophistication. A lot of times it's more like a known software bug or a weakness in a system that is known and people are looking essentially for someone who hasn't passed the system or updated the system. And that's another place where this is very difficult, especially if you have a big prop a sort of a big piece of infrastructure that you're managing is you have a lot of software that you just need to keep up the date, make sure people rotating passwords, make sure that you're rotating encryption keys, all this kind of stuff. And it just becomes a lot to manage and keep up the date. And so you're taking on a lot of responsibility there, it's kind of like back to the analogy that I was talking about at the beginning of, you know, if you store your money that you make in your mattress, you're taking on a lot of responsibility and you become potentially like an attack vector for somebody, you become a target because someone finds out that, oh, you know, Will's keeping all his money in his mattress, I'll just break into his house and take it. Yes, it's a two-fold problem. I mean, without getting too much into the details of what Skyflow does, but I'm intrigued, how do you build a product that deals with that? Because if my, yeah, if this date is all over the place and not necessarily in things that look like databases, like say log files and other things, then there's social engineering ways into get someone's password and all that. How do you put a fence around that? So I think the key is to not put a fence around it. I think that's a mistake because you're essentially, and that's historically what we've done is we've created essentially different point solutions. It's like, okay, well, we know we need to protect it here in the database where we're storing it. So we're going to add in some, you know, where of encryption, it will add in some policy controls so that we can make sure that not necessarily everybody has data. And then we have to do the same thing at the warehouse, and then we have to do the same thing in the log files, and H1 is essentially an independent different product. And then at some point, we'll need to decrypt the data so then how do we protect the servers that where we're decrypting the information because it allows them to play text in those. So you end up with all this path work. So I think that is the mistake. And what we need to do is actually take a fundamental look at the problem and take a first principles approach to solving this problem. And then a couple of companies that did this. So companies like Google, Netflix, Apple, Goldman Sachs, they took more of a first principles approach to this problem. And they were some of the pioneers of this concept of a data privacy vault, which essentially is a form of technology to specifically design for storing, protecting, giving you a situation of protection and governance over sensitive customer data. That's isolated outside of your existing system. So instead of having all these copies and your database and log files and so forth, you're creating essentially a single copy of that information that was within the vault. And then what you're doing is you're in all those different locations is you're storing kind of like a reference to it that's been anonymized. It's just kind of like a pointer. So that way, if you needed to delete will information because you said, hey, I want you, I don't want to be associated. I don't want you to have my information anymore. You only have to delete essentially that single source of truth. And then all those references essentially are meaningless. So you don't have to go find those thousands of different locations because those references become meaningless afterwards. So that is really the sort of best in class approach, but historically only really well funded big companies have developed a technology like that because it takes a lot of technical expertise to do something like that. If it's not your core product, not a customer facing features hired to sort of justify the cost of putting 50 engineers into something like that. If you look at Shopify, they have a blog post about this. They also presented this in the conference a few years ago. They did this for their analytics pipeline. It took them three years and contributions from I think 94 engineers, like that's a big expensive project for something that isn't necessarily like your core value as a business. So it's just not realistic for most companies to do this. So essentially that was like a lot of the inspiration behind Skyflow is like, hey, can we take this concept and bring it to the world that anybody can use? Hello, a quick reminder from me that if you're enjoying our podcast series, why not become a member of the DMI so that you can enjoy loads more content from webinars and case studies to toolkits and more real life insights from the world of digital marketing. Head to digital marketing institute dot com forward slash ahead of the game to sign up for free. Now back to the podcast. How do you how do you get people to care about that? How do you get your target customers to care? What's the kind of core value proposition in your marketing and how did you come up with that? Well, there's, I think there's two things. And in terms of getting people to care, I think that a big challenge that we have is that we are a net sort of new category of product. Not everybody knows what the data privacy vault is. It's a little bit like bringing bongo DB to the world back in 2011 when no one knew what a document, no sequel store was. And so you have to do a lot of awareness and education of the market, but you also have to educate the market on why should they care about this thing. At the same time as we talked about earlier, there is all this pressure. There is more and more pressure on companies to actually do something, especially as they're storing certain types of information or they want to operate their business in certain parts of the world. So GDPR of course is like forcing function for some companies. If you're storing credit card data, then you have to comply with PCI compliance regulations. And then if you're storing healthcare data in the United States, then you have HIPAA. So there are certain things in a certain highly regulated industries where even companies that are starting as startups or their established companies know they they have a problem they need to solve it. In terms of our original go-to-market, we really focus initially on sort of those payments like PCI use cases, thin tech space where they know even if you're a new company, I have to do something about this. And most of them aren't necessarily in a position where they're going to build out their own like PCI compliant infrastructure. They're going to look for some kind of vendor. So that was like a wedge into the market to start with because you could focus on people who have that particular problem. And then from there we expanded into a health tech as well where you have HIPAA compliance or the other challenges. People know that they're storing information about patients, but I need to enable my data scientists to be able to do their job. But how do I do that in a way where they are not essentially compromising the privacy of my patients records and so forth? So there are these sort of acute problems that we've been able to identify that people need to solve. And it's really about essentially building campaigns focused on those particular problems. So someone might not necessarily be searching for data privacy vault or even data privacy solution, but they do care about specific things. So if you're an engineer and you look on places like Stack Overflow, which is a very popular place for engineers to ask questions, you can find people asking questions about like how do I safely store social security number or how do I keep customer information out of my log files and stuff. So there are these problems that you can identify if you do some research that people actually care about. So the pain points with the kind of with technical teams and yeah, you think you're kind of intercepting them at just the right time when they're kind of researching that problem. It's interesting that yeah, on your website, you kind of the currently the head header is your tagline is what if privacy had an API? And you got this diagram showing, you know, an example of a customer record and how that's being accessed by things like HubSpot, but also OpenAI and other LLMs and customer data platforms and other apps like Stripe, Payment apps and things like that. That's that's that's very clever, like you say to this central secure source, but I'm just always interested in in, you know, in B2B how we sell quite, I don't know, things that people should have, but they don't kind of maybe put enough time into thinking about right. It's the kind of, you know, telling people, telling people about a problem they maybe don't fully understand that they have and trying to sell a solution for a problem they don't fully understand they have. Do you feel like you're doing that to a point or are people coming to you just at that point in their journey where they do know they need you. So I think they usually come, like I said, with a specific problem in mind and then sometimes there is, you know, part of our job is to educate them maybe on the larger picture or you sell them essentially, you know, our product is a platform so it's very flexible to solve holistically all these different problems that you might face as a business in the data privacy security space. So you might essentially alleviate this initial pain point, but then it's like a land and expand situation where it's like, okay, well, we could fix this particular problem for you. But then, you know, six months down the road now that you understand how to find it works like you probably have these other problems that we can help with and it kind of opens their eyes to the fact that like, oh, I don't need this just together like a hundred different products to solve this problem. So a good example of this is something like data residency, which is the problem where certain regulations privacy regulations in the world say that if you're going to essentially store or process information about customers where those are citizens of specific regions or countries, then that information needs to stay within that region or country. So some examples of this is like South Africa has has a regulation like this Australia, Germany, there's more and more of these regulation of parts of the world that say like, hey, if you're going to store information about our citizens, you need to store it in a data center within the bounds of our country. And then there's a bunch of regulations around how can you transfer data in and out and different places have different like levels of like strictness on this like China being one of the most strict places of the world for this. So that's a really hard problem for most companies because there's not really a good solution to that other than taking maybe your existing cloud center like maybe you're running a bunch of stuff on Amazon Web Services and basically copying it and redeploying it to another place. And so then you have you might have a data center essentially your entire cloud infrastructure running in Germany and you have one in the US and it's like, oh, well, we need to also do this in South Africa. Now I have another one in South Africa. It's hard enough to do this properly at scale in one region, let alone running like 10 regions simultaneously. It just becomes very expensive and a huge maintenance nightmare and then it introduces new problems where you have data siloed in different places. And now if I need to run analytics or do you know, build an ML model or do the science at the global level, I have no way to sort of consolidate all those different records. So that's a problem that is a big barrier to go to market for a lot of companies and we have a number of customers that come to us with that problem and we can solve it very simply by essentially just taking the regulated data and deploying vaults within those regions. And then suddenly you're enabled to run go to market there and because we're creating those anonymized references that you can store within your global data center, your sort of global data center is descopped from the client's regulation rate and agency. So when Nova is a good example of this, which is a customer of ours, where they needed to launch a laptop in like 20 different countries and they were using HubSpot for marketing automation for this, but HubSpot, you can only run HubSpot in Europe. In the United States and six of the countries that they were doing this and had data residency requirements. So how do you do that effectively and you can do that essentially by what we did was we deployed six different vaults and rather than putting the regulated in HubSpot, we just essentially stored the references within the HubSpot. So we descopped HubSpot from essentially the data residency requirement and it's suddenly now you're not restricted to only being able to sell to customers in Europe in the United States. You can sell customers anywhere in the world. Sorry, interesting that that is interesting. Yeah, no, I just I was interested, I'm interested about the marketing side of it as well because I just so you're not going out into the market with you're not doing like outbound advertising saying stop having breaches and sort of with these kind of pain points and fear points really. Yeah, we try to stay away from sort of the fear based marketing as much as possible, like I think people generally understand that, you know, data breach is bad. They might not necessarily know that there's ways to solve it, but I think there's ways that like I think position that message that's a little feels a little less like a scare tactic. I also think that it is be the wrong move to go to companies that just had a data breach like and try to scare them into using you because they're already in a place where they I think are very vulnerable. I think the better tactic is to have empathy for those people and understand that this is these are hard problems to solve and you know, maybe we can talk about it and figure out a way that we can help you, but we don't have this actually scare them into talking to us. So Skyflow allows you to pass customer data into GPT-4 for processing doing various clever smart things with that. But without actually revealing the customer's real information to GPT-4, how does that work? That's a good question. So a little bit depends on how you want to do it. That's essentially some of this is down the configuration, but as a simple example, instead of my name, Sean Faulkner, I can replace that with essentially name to essentially give enough context. It'd be like a prefix on the data that would say name, colon or something like that. So then GPT-4 in this case knows that this is someone's name, and then replacing the name with something like a UUID or some random string essentially. So name, colon, ABC123, and then every time Sean Faulkner's part of the input set, it would get replaced the same way. So I'm just consistently generating essentially a random string that is a stand-in for the original name. Because there's all that's vectorized in space, in high-dimensional space. Anyway, it's just an numeric value in high-dimensional space. As long as you don't need the LLM to have some language-based contextual understanding of the name, Sean, and all the things that means, or your address and where that actually is in the states. Do you know what I mean? So as long as you don't need it to know that thing, and you just need it to sort of identify that and identify patterns between the bits of data, you're good, right? Yeah, but even from a context standpoint, if you're doing this for all training data, then GPT-4, whatever you're using for your baseline LLM, your foundation model, you could still draw certain, or if you're building a model from scratch, even better. But you still be able to draw essentially relationships between those things, because I'm still going to de-identify United States the same way every time. So then I can actually, as long as I'm replacing the United States the same way, and that's part of an address, then you can still draw that relationship essentially between the representation of the state and the representation of a country and so forth. And the thing that these models are really, really great at is essentially deriving these types of rules and relationships without us explicitly telling them, we're basically telling them through the language that we create. The whole idea is for the best sort of approach to solving the challenge around data privacy for these models is to essentially never share sensitive data with the model. Instead, essentially give it a clean form of data, you're essentially keeping the sense of data out. And then when it comes to inference, which is the process of, I'm going to type in, you know, who's the president of a country or something like that, then that's going to run through the model, the model's going to create some sort of response. You can essentially place Skyflow ahead of that process as well. So I type in something like, who is Sean Faulkner, I'm going to pass that through Skyflow. Skyflow is going to, essentially, identify Sean Faulkner, remove it, replace it with the de-identify forms of data, and then the de-identify clean prompt is going to go to the model, and then the response will come back. The response has de-identifiers in it as well. We can automatically replace those. And essentially apply policies that you can are in control of to make sure that the person who's actually seeing the data has the rights to see it. And you can even control the format. That way, you know, will you get, you know, maybe you get a mass version of my name or a mass version of my email or phone number or something like that. But me is the owner of the information I can see all the raw data. And that is really the problem that is hard with these models is, how do you actually control, you know, who sees what when, where, for how long, and so forth. And there really is nothing else in market that allows you to have sort of that fine grain access control over the information in the model outputs. Yeah, well, fascinating that. Yeah. I mean, you'll know more about this than me, but it's sort of like the hashing of passwords that we've always done in secure systems in the background where the system never actually knows what the password is. It just gets a jumbled up hashed version of it and knows whether it's right or not. Yeah, exactly. That sounds very, very practical and useful. If you will get a history of passwords like 20 years ago, we used to put passwords, plain text and people's databases. And then, like, finally, we realized that was a bad idea. And essentially, what we figured out was that the sort of the only use case that you have for a password is to figure out does it exist or not. And you can solve that problem by essentially destroying the password through salting and hashing. And then you don't have to store the plain text version, which for the listeners benefit is turn the password through cryptography or something similar into a long string of meaningless characters. But it will always get turned into that with the same process. And so the system will just know whether the passwords correct. It just will never know in plain text that my password was. You know, brown dog 58 or something. Yeah, exactly. Yeah, it's basically irreversible. Yeah, it's like a irreversible. It's kind of like a random string, but it's essentially the same process each time. So if I'm a hashing, you know, my name, I'm going to get the same essentially output of the hash. So that fixed the password problem. And I think the key was there was recognizing that the only use case with a password was to essentially check to see whether it exists or not. Now what we've been able to do with Skyflow and part of our secret sauce was we applied kind of that same way of thinking, but to all customer data. So if you look at something like a credit card number, for example, then a credit card number there is very specific things that you want to do with a credit card. Like there's only a couple of like legitimate use cases of a credit card. And essentially that dictates how if you think about it at that level, that determines how you need to store that information so that you can support essentially something like fully encrypted operations, which means you can live in a fully encrypted work. Similarly, the problem has been that, you know, we can encrypt data, but then in order to make it usable because encryption basically breaks a lot of systems like search, you need to decrypt it. And then if you decrypt it, then you're vulnerable at that point because someone could essentially get access to that information. What we've been able to figure out is a way to essentially stay in the world of encryption, like keep everything encrypted, but satisfy these different use cases. The same way that you're doing this with a password where you're hashing it, but like a credit card, the only use cases are, I need the last four digits for verification. And then if I need to pass it to a payment service provider to carry out a transaction against a credit card, then I need to secure a way to essentially pass the full credit card. It's a way to check to what and see whether it exists. And essentially every form of like customer data or PII has these specific use cases for it. And we've been able to essentially apply, like think through all those different use cases so that we can essentially store the data in a special way, which we call polymorphic data encryption, which allows you to essentially keep this day in this fully encrypted world. And the trick there is that to stop thinking about these things as something like a number or a string, like a credit card, we call it a credit card number or password number. They're not really numbers. They're like actual data structures. Like you don't take a credit card number and multiply it by a passport number and divide it by a social security number. Like that's not a real operation. The real operations are essentially, I need to show only part of the information or I need to check whether it exists or I need to pass it to a third party system in a secure way. That's pretty much it. So if you think through all those different use cases for the hundreds of different PII, essentially you can create a very secure system that allows you to perform all these operations fully encrypted. Interesting. I feel like you've probably been asked this question before and maybe by your friends and colleagues you get asked this. What do you say to people who ask you then about the safety of the information that I put into an LLM like chat GPT? So what do you perceive as being the risks for me? I'm working at a tech company. I put a bit of code in just to have it check a bit of code or I work in a e-commerce company and I took some customer records in to help. So it can sort of reorder them and maybe tidy them up for me. What is the actual risk with that? Do you think? So I think there's a couple things. Originally it wasn't clear from in the chat GPT world what the data was actually being used for. If I essentially put in a prompt that has like a credit card information, what actually happens that information? Like how long do they hold on to it? Are they using it later for improving the model and so forth? I think they've cleared up and tried to clarify some of those things that the information is deleted after some point of time. But even if you're using the APIs from an AI, by default, originally a lot of the stuff that you were doing from inference was actually used for training later and it was essentially something that you had to turn off. So there's all these little little things that you have to be conscious of. But even if they're only holding on to the data for a certain period of time with chat GPT, they still have that information somewhere in their systems. And so then you have to trust, are they doing a good job of protecting it or not. But the other big problem is that despite how well-resourced a company like OpenAI is and a pick on them for a moment and how much probably time and attention they're putting to the problem of like ethics and privacy and security, they still people have figured out ways of exploiting the system. Recently some Google researchers were able to prompt engineer chat GPT to give up a bunch of the raw data that was used during training that actually had customer PII or people's PII in it. They did this through getting it to repeat the word poem infinitely and then eventually it started actually spitting out the training data. If a company as well-resourced as OpenAI and putting all this time and attention to it can't handle this problem because I think a part of it is all so new, it's hard to know exactly where these problems could occur. If they can't solve this problem, it's very unlikely that if you're like a three person startup that are investing in these technologies that you're going to be able to figure it out. So I think that's just something that we all need to be conscious of and be asking ourselves these hard questions. Are we in a position where we can actually address this issue? No, that's true. That's a good answer. It's like all these things. It's unlikely that there's ever going to be a problem. It doesn't sound great. Does it send in your confidential data into a black box and just hoping that the people who own that black box keep it locked? Yeah, exactly. I think because of the modality that you're interacting with something like ChatGPT feels somewhat more personal because it's not like, I think we've gotten to a place with something like Google Search where we're probably not going to put in our social security number or credit card number or something like that because we've been trained for 20 years how to essentially query information. And generally, we're doing it short, three words snippets and so forth. But when you interact with something like ChatGPT, it feels you can finally ask a computer a question like you would ask a normal human being and then you're getting something that feels like a human response. And I think that kind of changes the level of comfort that you have as someone like using the product where you might feel more comfortable sort of putting certain things in. It's very easy to do even even outside of that where I think a very natural thing to do is take a contract or take some wrong form document and say summarize this information for me. And then that might have certain sense of the data in it. Like there was a recently a lawyer that gotten some trouble with doing some stuff where they were crafting contracts through ChatGPT and so forth. So like all these types of things like the efficiency gains that you want to get out of the product become a place where people might not even be thinking about the potential risk of what could happen. But there are ways of solving these problems like recently Amazon just before AWS reinvent their big tech conference at the end of the year. It took place last month. They announced a product called AWS Party Rock and what it is is an LLM playground where anybody without any technical skills can essentially create an LLM powered application just through writing like a prompt to be like, hey, I want to create an app that will help me figure out an agenda for a podcast or something like that. And it'll just make designing UI for you in this such way you can start playing within take advantage of these things. And there is another example where I could easily paste in contract information and ask if the summarize like I could have some kind of contract, you know, summarization tool or Q&A tool and so forth. But what I was able to do ahead of that event was I actually took Skyflow technology and I created was a Chrome extension which is such we can install into Google Chrome and it would monitor any of the inputs that you put in the party rock and then run that through Skyflow and essentially replace any sense of data with the identify forms of data. And then when you get the response back from party, but rock, it would re-identify so you can use it essentially as like this firewall on top of a system like that and party rock still functions exactly as you would expect, but you're essentially buffering the risk of putting personal information into it. That's very interesting. Yeah, I must give it a go. What do you as a marketer, what have you seen the real impact of AI on just on your work? I mean, you talked about it being a copilot and a kind of brainstorming buddy and what have you, but amongst your staff, has it been the game changer that people talk about it being. So I think right now we're still like very early, like I would say that we're kind of in the like pets calm era of the AI world, like we are sort of I think moving from the non AI era into an AI era. So this is I think like something that's going to have major impact, but I think in terms of feeling real impact, if like significant impact to people's companies and maybe even places where you know certain jobs go away, I think we're still early on that. I think the biggest thing that people are doing right now is kind of this like copilot chatbot experience was kind of like the like base level of using this technology is like the natural thing is like, hey, we'll stick a chip chatbot in the application or something like that and people can talk to it in a new way. And those are useful and there I think they also can help you be more efficient it you know, it certainly helps solve like the the blank page problem that people will face sometimes when it comes to you know writing something and that's very useful and helpful, but I think that the long term impact is going to be much bigger like there hasn't been essentially anybody because still too early. That really has built like we thought something like marketing automation or a CRM from like an AI first lens and I think that's something in other you know marketing tools and so forth like there's I think in the next five to 10 years we're going to see a whole sort of net new series of products that are applying you know AI from first principles that's just baked in the product that where they're much more adaptable. They adapt to your needs and so forth and they understand potentially patterns of data at a much deeper level rather than I think like a lot of times like as a marketer right now, especially if you're like a mark ops role or you're in like growth or something that you're a lot of times you're having to go out my proactively pull data run reports and so forth and digging the stuff at some point. AI systems are going to do a lot of that work for you and actually surface potential places where you can make improvements and so forth so I think that is the sort of the next stage of this but we're probably a few years away from actually seeing them yeah it you're right it does feel that way for sure. It's sort of a bolt on it's just getting sort of bolted on to everything and exactly yeah and not not always in a massively useful way and just I'm interested in your trajectory as well like your I mean your head of marketing its sky flow but you renal you come from a technical background. So developer relations in places like Google how and you founded a tech company as well how did you wind up in marketing or what drew you to come out of marketing yeah I guess I'm in some ways like an accident I'm like I my I spent 10 years studying computer science so I was on a path of being a professor or professional researcher some form and I was doing a postdoc in at the time of bioinformatics. Accessor completing my PhD in computer science and that's when I actually ended up starting a company with a couple other students and then we raised some capital and I left the world of academics to go and build this company and I was the CTO of that company so I was the technical co founder so I built most of the original product did nearly you know all the engineering work for the first year and so forth and then ran like the engineering team but through the process of actually building a company. We were never super you know heavily resources so we always had to figure stuff out our own including essentially marketing and sales and so forth so those were things I didn't really have any experience with I was very comfortable with the technology side but I had to learn because there's basically no one else to do this thing so I built like our marketing team and I learned a lot about content marketing and SEO and so forth just because I knew that we needed to figure out a way to do go to market at scale for a low price product and do like a PLG motion and so forth but if I didn't figure it out we were going to fail as a company so it's like nothing you know focuses you to essentially like the prospect of running at a capital or running at a money to pay your employees so that really became a forcing function and then you know I helped build our SDR or SDR organization and so forth so I learned a lot about sales and marketing through that experience and then from there I you know I ran the company for seven years I'm going to try to figure out what I was going to do next after we we sold the company and I felt like originally I was going to go back to being an engineer because I was like you know I spent a decade studying this stuff I should probably leverage the skills more on a day-to-day basis but as I started to explore roles I was a little bit concerned about getting bored with doing pure engineering because I was so used to wearing like 50 different hats so I was originally referred to Google as a software engineer but then when they saw my background as like somebody taught at university and blogged and written for a long time presented a conference is started a company they approach they said like hey we think you'd be a great fit for development relations is that something you're interested in I was like awesome you know what is development relations because I had no idea what this thing was and then they explained the different roles within Google and the one that really appealed to me at the time was a developer advocate which is now known as a developer relations engineer at Google but it combined sort of being an engineer but also an educator and some of the sort of marketing skills I had developed as well. You might be creating content and you're a little bit more in control sort of like developer go to market and then I had a really exciting opportunity at Google where I even though I was joining a big company I got to be sort of the founder of the developer relations team for a net new project there. So I you know it's actually own the entire developer go to market developer relations development experience and by the time I left I built a team and we were doing that for four different products and then I joined Skyflow originally as head of the relations I just love the vision of the product the leadership team and on to Sharma the CEO and co founder was an investor in the company I started so we had known each other for about a decade. And I originally joined to do such way the same thing that I've done at Google but through just you know a series of things that happen within the company you know after four or five months I took over product marketing as well and then about a year ago I ended up leading all of marketing so it's been in a fantastic experience so far so far and I really enjoyed it and I think the key with some of the stuff like I don't pretend to know everything that there is about marketing but the key. Yeah exactly like the key is low ego higher people that fill in the gaps for you and that are really good and give them the space to basically do their job effectively and I think the thing that I have that is makes me maybe the the right person for the job company we have is that because I was a CTO which is essentially our core buyer persona. I understand how they think about the world how we need the message things how we need to position how do you reach those types of people so I'm in a unique position where I have marketing experience but I also have experience as our core buyer persona and I think that I understand the product in a deep level and I have good instincts essentially when it comes to. Like figuring out how do I actually you know meet these people at the place that we need to meet them that that's interesting that what I'll ask you what you don't like so much but what do you love most about working in marketing so I guess like one big thing for me is that I feel like there's always something new to learn and I really thrive on sort of the I'm like pushing myself to get better and like learn as long as I'm learning I feel really engaged and happy and there's like you said you can't know everything. You can't know everything in market is such a breath like you know we kind of as an outsider you you think of this as maybe like a monolith but there's so much stuff within marketing from your PR and comms to digital ads and growth to even sometimes develop relations of all for marketing like there's a huge breadth of essentially functional areas that are fun to like dive into and gain experience and. I mean it's simultaneously one of the best things it's definitely one of the things that drew me to marketing. I also can be frustrating as well because there's days where you just feel like there's so many things you could be doing maybe should be doing and there's just never enough hands on deck and that I think that drives there's always a new shiny ball to get your head around and that can drive some people mad and I think drive some people to burn out as well. Yeah absolutely. I think that's something that being a founder has like helped because you know companies generally I think I always say like a startup is more likely to die from indigestion of a surplus of too many good ideas than starvation of too few good ideas. So when you're the leading a company a big part of like it's such a focus is such a like precious thing so you you either figure out how do I focus on the right things like what are the you know free right good ideas that I should be focused on and prioritize those and test and iterate or you're just you know you're going to die as a company and you know I don't think I have certainly always a hundred percent success rate with with with maintaining that level of focus because like you said. There's always comes you know new shiny ball. There's lots of requests that come in and you want to make people happy and so forth but a big part of the role of any leader with an organization is figuring out when is it appropriate time to say no and push back on things and keep people focused and I think that's really the key to being successful. Yeah and you've led a few teams I mean now you're head of a marketing team. How do you keep your staff saying how do you how do you help guide them towards not burning out in particularly in marketing but also I'm sure that's the case in in tech as well. So I mean I think that there's a few different things like a lot of it comes down to the culture of the company and a lot of that comes from the leadership team. I think that one of the things that attracted me to sky from from the beginning is that the leadership team across the board are very experienced people. They've all had stints in big tech companies as well as time in startups so they have a good balance and they understand if you want to build a big company you want to build a snowflake size company or Salesforce size company you know it's a marathon. You can't just drive people into the ground working 80 hour weeks and so forth like certainly there's times when you have to push like in any company but you can't just drive people into the ground like you're just going to burn people out and and and not lead to long term success. I think that's something that you have to instill from from the top down and that feeds into you know how you hire how you you know prioritize things how you encourage people what do you reward as well you know and and I think we've always had a culture where it's okay to take a break from work and I try to give people that space on the team. I think that is a challenging time in the market of the tech world because there has been you know construction in a market it's not as easy to raise capital as it was a couple of years ago stuff and I feel very you know privileged and happy to be in the position I am at Skyflow there's a number of companies that I talk to when I was considering leaving Google that I probably wouldn't be at most companies anymore because they've had massive layoffs so we've been very smart and conservative in terms of how we've hired and tied those. To revenue goals and so forth and I think a lot of that's you know comes down to the experience of the leadership team yeah it's good yeah definitely mature leadership is always a good thing thinking about our audience of digital marketers what your key tips to when when they're thinking about better practice around their customer data I think the key is that we tend to optimize the things that we measure so you need to be careful about what you measure. So whenever you're essentially you know putting for some sort of KPI or metric that you're communicating you're trying to optimize more you need to be asking I think deep questions about is this the right thing to be measuring what does this say because there's always trade offs you know if I'm you know in a demand general and I'm optimizing for mqls then there could be trade offs in terms of the quality of those things like you know I might be able to generate lots of you know contact form fills but I can also do that by putting a free beer. So I need to be asking these questions of like what does this actually say what is you know quality and look at those things downstream and figure out like every you know quarter so is this the right thing that I should be optimizing for the right thing that I should be measuring. Yeah that's that's that's a good tip what do you think 2024 holds I don't need a full trends prediction don't worry what do you think what do you think 2024 holds in the sort of world of certainly this intersection of privacy and technology. So I think that we're going to see a lot more regulations starting to be pushed out with AI there's already works in the EU with the AI EU AI regulation act president Biden came up with an executive order in the United States earlier this year where a lot of it was about concerns around privacy for AI systems. And I think we're in this like huge hype cycle so I think what I predict for 2024 is we're going to start to go in through come out of the hype cycle and go into the profit disillusionment and then the real work is going to start and I think I'm looking forward to that you know I think it'd be great to have a little less noise in my news feed about all the things going on and generally AI and actually start to see some like. Real real projects land that I have meaningful impact to people's lives same here I'm very much looking forward to that I think good stuff well thanks so much I really I feel like I've learned an awful light in the last hour that's been really good honestly I really appreciate your time well only one quick question to ask you where can people connect with you and find out more about you online. Waitins probably the best place I'm also of course on Twitter but if you just search my name Sean Faulkner I should come up on both places and if you want to check out Skyflow you can check us out at skyflow.com. We will do that I'm definitely going to have a play with it and look thanks so much Sean again really appreciate your time thanks very much. Thank you so much cheers. Thanks for listening.

Podcast Summary

Key Points:

  1. Skyflow provides a data privacy vault service, offering a centralized, secure solution for protecting sensitive customer data, inspired by approaches used by major tech companies.
  2. Data breaches are increasing despite significant corporate spending on security, partly because data protection is often not a core business focus, leading to fragmented and vulnerable data storage.
  3. The complexity of modern data systems, with information replicated across numerous locations (databases, logs, backups), makes comprehensive protection extremely difficult using traditional, piecemeal security tools.
  4. Common breach methods include social engineering and exploiting known software vulnerabilities, with AI potentially increasing the sophistication of attacks.
  5. Skyflow's go-to-market strategy initially targets regulated industries like payments (PCI compliance) and healthcare (HIPAA), where data protection is a mandatory and acute business need.

Summary:

This podcast episode features Sean Falconer from Skyflow discussing the growing challenges of data privacy and AI. He explains that despite massive investments, data breaches are rising because companies often treat data protection as a secondary concern, not a core business function. This leads to sensitive customer information being scattered and replicated across thousands of systems, creating an unmanageable attack surface that traditional security tools cannot adequately protect.

Skyflow's solution is a data privacy vault—a centralized, secure repository for sensitive data. Instead of storing actual personal information in multiple places, companies store anonymized references or pointers to the single, protected copy within the vault. This simplifies governance, reduces risk, and is based on a model pioneered by large tech firms.

Skyflow initially markets to industries with strict regulatory requirements, such as finance and healthcare, where the need for robust data privacy solutions is already well-understood and critical for compliance.

FAQs

Skyflow provides a data privacy vault service designed to protect sensitive customer information. It addresses the challenge of data breaches by offering a centralized, secure solution instead of scattered data storage.

Many companies treat data privacy as a secondary concern, focusing resources on customer-facing features instead. Additionally, legacy systems and scattered data storage make it difficult to secure information effectively.

PII refers to sensitive customer information that can identify an individual, such as addresses, phone numbers, or social security numbers. It includes both regulated data and other details companies want to protect from leaks.

Breaches often happen through social engineering, known software bugs, or unpatched system weaknesses. Attackers may exploit human error or outdated security measures to gain unauthorized access.

A data privacy vault centralizes sensitive data in one secure location, replacing scattered copies with anonymized references. This simplifies protection, governance, and deletion requests, reducing breach risks.

Large companies like Google or Netflix have built vaults, but they require significant resources. Most businesses lack the expertise or budget to develop such solutions internally, making them impractical.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.