Go back

#37 - Identity Categorization Done Right

35m 40s

#37 - Identity Categorization Done Right

The episode challenges the common practice of treating identity as a static, labeled noun in identity and access management (IAM). It introduces the concept of the "uncanny valley" — where highly detailed taxonomies look mature but fail to translate into actual differences in security controls. The core argument is that identity behaves like a verb, not a record: it's defined by its actions, context, and risk over time. Instead of relying on complex labels like RPA bots or workload identities, organizations should focus on three behavioral dimensions: interactivity (human vs. machine), lifespan (ephemeral vs. persistent), and privilege (sensitive access vs. standard). Each combination creates a distinct, governable behavior pattern. The episode emphasizes that taxonomic depth without corresponding control enforcement results in "taxonomic debt"—a gap between label complexity and real security. A healthier approach is to design only as many categories as can be backed by enforceable policies, then continuously refine based on telemetry and behavior. The ultimate goal is to move from label-based governance to behavior-based automation: detecting anomalies like a non-interactive identity logging in interactively or a temporary identity lasting too long. This shift enables real risk reduction, improves visibility, and prevents security theater. The challenge to practitioners is to audit their current categories and identify which ones lack real behavioral differences—merging or retiring those that are merely cosmetic. By anchoring identity governance on behavior, organizations create resilient, responsive, and operationally sound systems, avoiding the illusion of maturity that comes from fancy diagrams.

Transcription

4301 Words, 25331 Characters

English
Alright, story time. We could be all working for any big bang, a tech company, and healthcare company, or whatever. But I am sure we have all seen that absolutely gorgeous I am slight. You know the one with boxes, colors, arrows, and 15 type of identity accounts. It looks like the marvel multiverse of identities. And for a second, we are all impressed. But every time I see that slide now, my brain does this little tilt, and I find myself asking the same question every time, okay. But does any of this actually change what these identities are allowed to do? Welcome back to the identity navigator. I am Rohit. And today we are going to hang out in this weird little corner of identity and access management that I am calling the uncanny valley of identity categorization. You know the uncanny valley, right? The feeling when a CGI character or a robot or a picture looks almost human, but not quite. And your brain goes, nope, something is off. A lot of I am taxonomies give me that same vibe now. They look mature, they sound mature. But when you get close and actually trace that controls, they are unsettling. So in this episode, I want to unpack a simple and slightly uncomfortable idea. We treat identity as a noun, a static record in a directory. But in any real enterprise, identity behave like a work. A continuous state of risk, context and relationship. And the big mistake I keep seeing is not, you don't have the right tools. It's something I call taxonomic depth, where your categories evolve faster than your controls. If that sounds abstract, don't worry. I will tell some stories, we'll walk through some examples and by the end, you'll probably see pieces of your own environment staring back at you. So let's ground this in reality for a second. Image a report you look at now basically says the same thing. Non-human and machine identities massively outnumber human identities in enterprises. In some environments, it tends to even hundreds to one, which is correct. So if you picture your employee directory as a room with say a thousand people in it, your actual identity landscape is a stadium packed with machines and a little VIP balcony with the humans. The problem is, a lot of our I am thinking it's still human first. We obsess over employees versus contractors versus vendors versus partners versus intern. While on the machine side, we just keep inventing more and more labels and kind of hopes that governance. So you end up with taxonomies like RPA, but service account, system account, workload identity, integration user, AI agent and everyone feels good because hey, look at all these nuance. Surely with this many categories, we must be doing some serious governance. But here is where the uncanny valley creeps in. Those labels look sophisticated until you ask. School bro, but what's actually different in the control plane for each one of those. So let's zoom way out for a second and talk philosophy. Most I am systems treat identity like a record. It's a row in a database and object in a directory or adjacent blob in some identity store. This name, last name, employee ID, department and maybe a couple of flags. Is this a service account? Is this an RPA bot? Is this a workload identity? It's all very noun like it just sits there. But if you look at how risk actually shows up in real enterprise, identity isn't a row. It's a pattern of behavior over time. Think about it, a human identity that never logs in is boring until one day it logs in from a new country at 3 a.m. and immediately goes for a high value app. A service principle that normally calls one API in one region is fine until it suddenly fans out across regions and start touching data it never needed for. A temporary workload identity that was supposed to be short lived quietly becomes immortal because nobody wired in rotation or revocation. Those are verbs that's identities as doing not just identity as being. And here is the subtle trap when you think of identity as a noun, adding categories and labels feels like progress. When you think of identity as a verb, categories are only useful if they drive different enforced behavior in your control plane. That gap between how fancy your taxonomy is and how little it changes behavior, that's where taxonomic depth lives. So let's put a name on this. The basic depth is what you accumulate when your identity categories evolve faster than your ability to enforce distinct controls for each category. So it sounds like this, we distinguish between RPA bots, service accounts, system accounts, workload identities, vendor technical users and AI agents and each one has a policy document and a naming convention and everyone in that meeting feels just a little bit better. You have got a model, you have done governance, obviously since you are naming everything you must know what you are talking about. But here is a litmus test, I want you to try the next time you look at your own taxonomy. For each category, can you point to at least one real enforced control difference compared to its neighbor? Now you cannot cheat, you cannot tell me, you cannot point to a bullet in a policy doc, not a promise in a deck, but a real control. That is what you have to point to like different authentication requirement or different credentials lifetime or rotation rules or different monitoring and alerting thresholds or maybe different joiner mover lever behavior. Now if for two items or more than two items, if the answer is not really or yeah, we meant to but then that category is basically I am cosplay, you have created complexity without adding control, you have made the diagram better, you may be impressed few folks but you haven't made the environment better. That is taxonomic debt and that technical debt, the interest compounds, every new tool you add brings its own labels which you now need to map into your already complex schema. Every exception gets its own special subtype, your automation has to branch on 30 different account types that all behaves almost if not exactly the same. On the surface, it looks like you have got a level 5 maturity model and underneath you are still a level 2 enforcement just with fancy words. Now all of us in identity, we have to ensure that we are not driven by buzz words but actual operational maturity and this is what I am talking about here. So let me paint a very real word picture. Imagine an organization decides we are going to get serious about non-human identities. So they kick off a working group, they do workshops and they land on this very sophisticated taxonomy, RP bots, service accounts, system accounts, work load identities, technical integration user and each one gets a naming convention, a two page policy word document and a box on the architecture slide. You fast forward six months and you ask a very unsexy question, okay show me the actual difference in the control plane and you discover they are all in the same directory or cloud identity store where they all authenticate with long lived secrets or keys. None of them are enforced to use short lived tokens. [BLANK_AUDIO] of them have truly enforced rotation, its best effort. And monitoring is basically, is the account disabled, yes or no. So on paper, this environment looks mature. In reality, if an attacker compromises any one of those labels, the blast radius and the different posture are basically identical. The labels are different. The outcomes are the same. We have made better costumes for our identities, but we haven't made better decisions about what they can do. So now let's talk about the part that gets a little spicy, incentives. IM is not just about technology. It's a giant game of incentives and shortcuts. Every time you create a category that has less friction or fewer prompt or fewer checks, you have basically created a loophole with a marketing name. Here is a pattern I've seen so many times I could probably whiteboard it in my sleep. You create a service account category that is exempt from MFA. A void certain conditional access policies is allowed to be shared by multiple processes and has strongly recommended rotation, but nothing actually breaks if you don't rotate. The justification is always reasonable. It's a machine robot. We can't send an MFA push to a bot and that's totally fair on the surface. But the moment you do that, you haven't just categorized the machine. You have now created a low friction identity track inside your system. Now imagine you are a dev or an ops engineer under pressure to ship a new integration by Friday. On the right path, you have per user accounts with MFA approvals and maybe some awkward UX. On the service account path, you have one credential, no MFA, fewer policy checks and it just works. Where do you think people route over time? What happens is people start using these non-human accounts for human interactive tasks. They log into consoles, tools, even desktops using service accounts. You get shared credentials all over the place and audit trails becomes meaningless because a workload identity is actually a shared human backdoor. The label says machine, the behavior is extremely human. From a game theory perspective and all of you must have heard or some of you must have heard me talk about game theory all the time. This is totally predictable. If the easiest way to avoid friction is to pretend to be a machine, humans will absolutely pretend to be machines. Not because they are bad people, because they are rational people trying to get their job done in a system that accidentally rewarded the wrong behavior. So when you look at your categories, don't just ask, is this technically accurate? Ask what behavior does this label invite? If service accounts quietly translates to MFA free accounts with broad access, you have created an identity honeypot. But for your own teams and eventually for attackers. So let's get a bit nerdy for a second, which comes very natural to me. In I am an identity is basically just a pointer. It doesn't contain the human or the machine. It points to a set of credentials or secrets, a set of entitlements, a set of relationships like groups, roles and attributes and a history of behavior over time. All our beautiful categories of RPA bought system accounts, workload identities, those are just metadata hanging off that pointer. And here is the nuance that trips people's up. If your metadata is too specific, you shatter your visibility. If your metadata is too generate, you miss apply controls and burn your teams. So let's walk both sides. To specific, you go all in on taxonomy, you invent a subtype for every scenario. RPA bought finance, RPA bought HR, system account batch, system account real time, workload identity lambda, workload identity container, workload identity, vendor managed. You get the point. And believe me with this approach, your dashboards will look incredible. But the problem is you can't easily answer. Show me all non interactive identities with standing access to production data. Because that pattern is scattered across 15 different subtypes. Your policies are copy-pasted version of each other with tiny unexplained variations. Everything becomes a case statement from hell where every new subtype is another branch to maintain. So the risk isn't reduced. It just hidden in the metadata of 50 different labels. Now, let's walk the other side, which was too generic. On the other extreme, you say, no, this is too messy. I had a podcast once and we will just have to users and service accounts. Now your I am engine starts applying human controls, MFA prompts, password rotation, rules, help desk close to machine identities that are supposed to run unattended at 3 AM. What happens? Jaw fails because they can't pass an MFA challenge. People hard code secrets or bypass controls just to keep the lights on. Your team gets branded as the department of no and shadow I am pops up. So the lesson is not more labels good or fewer labels good or more labels bad or few labels bad. The real lesson is your taxonomy should only be as specific as your control plane can genuinely support every extra label that doesn't map to a real behavior is basically another line of taxonomic debt. So, it's not RPA versus system versus workload. What do we categorize on? Right. I've been talking about the problem for the longest time, which as we all know is a easy thing to do, but what's about the solution? So, here is the pivot I want you all to try. Stop anchoring on what the identity claims to be. Start anchoring on how it behaves. Instead of asking is this a service account or an RPA bought? Ask three much simpler questions. That's it. Is it interactive? Is it ephemeral? Is it privileged? So, let's walk through each but in a very concrete manner. Is it interactive? Interactive identities are one where there is a human behind a keyboard or a device. Someone starts a session, they click around, they type commands. Non-interactive identities are used by machines to call machines. API, demons, agents, bad jobs, workloads, no human is sitting there pressing approve every 30 seconds. The control story should be a night and day different. Interactive, strong, phishing, resistant MFA, device posture checks, session controls, may be continuous risk evaluation. Non-interactive, no human MFA, but strict credential lifecycle, scope tokens, perk work load identities and behavioral monitoring. So, your first behavioral split is interactive actor versus non-interactive actor. And here is the key rule. If an identity is labeled as non-interactive and you ever see it show up in an interactive flow like a web login or a console login or a VPN, that's not whoops we will fix it later. That's either a policy population or the beginning of an incident and you should be asking some hard questions. About not just who used it, but also about how did you allow it to be used in that manner. The second is, is it ephemeral? So, this is the second axis, the lifetime. Some identities are supposed to live for years. Employees, hopefully, long-term partners, core infra and others should have the lifespan of a snapchat message, spin up, do one thing, disappear. If you treat ephemeral identities like long-lived ones, you get stale credentials everywhere. Zombie permissions and the huge graveyard of a temporary access that quietly becomes permanent. If you treat long-lived identities like ephemeral ones, you drown in noise and break critical stuff. So we define ephemeral actors versus persistent actor. Controls then looks like ephemeral, born with an expiry, created and destroyed by automation, short-lived tokens, rotation as part of natural life cycle, and persistent, which is more stable identifiers with strong life cycle governance, tight-win org change, role change, and off-boarding. Now combine the first two axes and you already get something powerful. Interactive plus persistent is classic human user. Interactive plus ephemeral is just in time elevated sessions, break-glass access. Non-interactive plus persistent, long-lived system agents, and non-interactive plus ephemeral modern workload identities in cloud native architecture. And each of these deserves a different control recipe. Now let's look at our third axis privilege. This is the one where we are the most stuck in old mental models. A lot of organizations still behave as if only human accounts are totally privileged, even though many preaches now leverage over privileged machine identities and secrets. So ask this. Does this identity have standing access to sensitive data or admin functions? And if this thing gets popped, do we suddenly have a really bad day? If yes, it's privileged. I don't care if it is service account or RPA bought or AI agent, the behavior and the blast radius are what matters. Privileged actors, human or machines should be governed by just in time elevation whenever or wherever you can. Strong isolation of credentials and secrets. Enhanced monitoring and anomaly detection and foster revocation and incident response paths. When you map privilege onto the first two axis, you get eight interesting buckets. Interactive persistent privileged, interactive persistent non-privileged, interactive, interactive ephemeral privileged, interactive ephemeral non-privileged, you get the point. These are three combinations, eight categories. And now your taxonomy is not about marketing labels. Don't get fooled by the hype. Don't get fooled by the beautification of that slide. Don't get fooled by the nuances. Don't get fooled by the vendor slideware. Do what is right for your business. It's about how this identity behaves and how dangerous it is when it misbehaves. And for each of those buckets, you can design a concrete enforceable control plane. So let's connect this back to where identity and access management is headed. We are not going to win this game with a slightly better OU structure or one more level of directory testing. That era is gone. When non-human identities outnumbered humans by tens or hundreds to one, you physically cannot govern them at human speed with manual reviews and pretty diagrams. The future looks a lot more like this. You define a few meaningful behavioral categories, interactive versus non-interactive, a female versus persistent, privileged versus non-privileged. You attach real control recipes to those categories. You wire your telemetry to continuous ask is this identity behaving like it's category says it should. And just a footnote there. I'll keep AI agents out of that because that's an evolving field and I have more thoughts, but I did not wanted to muddy the waters here. So I'm specifically asking you to keep those out of here. But you can absolutely apply this model to them. I'm just saying there is a little bit more nuance and it comes to AI agents. So now if a supposedly non-interactive identity starts logging into web UIs, that's weird. If an FM reliability never expires, that's weird. And if a non-privileged identity suddenly starts doing privileged things, you guys did. That's weird. And when something is weird, the system doesn't just shrug. It reacts. Step-up challenges, revocations, quarantining sessions, paging humans. That's identity as a verb. The labels become hypothesis about behavior, not decorations in the directory. All right, let's poke one more sacred cow. If you build a level 5 taxonomy on top of a level 2 program, you don't get level 5 security. You get governance death. What does that actually look like? A level 5 taxonomy is extremely detailed diagrams. Lots of categories and subcategories. Beautiful architectural docs and everyone leaves the workshop feeling very smart. A level 2 program is some automation but lots of manual steps. In consistent ownership, partial coverage of apps and clouds and reviews that happens because audit is coming, not because you want them. When you attach a very advanced, very nuanced taxonomy to a very or a pretty basic implementation, you create fragility. Every new exception becomes a new category. Every new tool or SAS apps comes with label, you try to force fit into your model and nobody can hold the whole thing in their head so people start improvising. It's like bolting a formula one steering wheel onto a Karola. You don't suddenly go 200 miles an hour, you just make it harder for normal people to drive the car. A healthier approach, design only as many categories as you can back with real enforcement. Forget my access, come up with your own but you will have to have to back them with real enforcement. Let your taxonomy grow as your automation and telemetry mature. Be ruthless about retiring categories that turned out to be just labels. Every category is a promise. When you see this label, it means this identity will be governed differently. And if you can't keep that promise, that label is just dead. Right, Monday morning, how do we start untangling this? So what do you do with all this? Let's make this very practical. If you're listening to this on your commute or your walk, think of this as a little checklist you can drive with your team. First, grab a whiteboard or a dock and list out all the identity categories you have today. Human employees, contractors, interns, vendors, etc. Non-human service accounts, workload identities, RPA bots, technical accounts, integration users, whatever your word looks like. For each one, ask a brutally simple question. What are the enforced control differences for this category? If you cannot name at least one real behavior difference or rotation, monitoring, life cycle, that category is a consolidation candidate. This exercise alone will flush out a surprising amount of governance theater. Next, take either identity directly or your existing categories and tag them along the three access that we discussed. Ractor versus non-interactive, a funeral versus persistent, privileged versus non-privileged. Don't overthink it on day one, start with rough rules and refine later. The goal is to build a behavioral map of your landscape that cuts across vendors and legacy labeling. And you will immediately surface weirdness like, wait, why is this service account being used interactively? Or why is this temporary workload identity still around a year later? Or why does this non-privileged account have admin rules? Because these answers are gold. this is where your risk is hiding. Now for each meaningful combination, a interactive persistent privilege, write down what must be true about how you govern it. How is it created? How does it authenticate? How does it get access? And how is that access review? What's monitored by default? What happens when something looks off? Think of these as recipes, not policies. If you picked up this recipe and applied it in Octa, in Entra, in AWS, and in your SaaS applications, it should basically still make sense. That is when you know you are designing for behavior, not for one vendor's idea of an object type. Now once you have these behavioral recipes, go back to your taxonomy and ask category by category. Does this label map to a distinct recipe or it just a synonym for something you already have? If it is just a synonym, merge it. If it doesn't map to anything real, retire it. Yes, this might or will probably will annoy some people who are emotionally attached to their favorite acronym. I promise you, you will survive and your program will be easier to reason about. Now, last steps and this is where it gets fun. Start wiring your signals and automations to the behavioral model, not just to the labels in your directory. Use logs, context, and events to infer. Is this identity being used interactively or non-interactively? Is it behaving like something ephemeral or something persistent? Is it touching privileged paths or not? Then build responses around deviation, non-interactive identities, doing interactive stuff, flag it, or whatever else you want to do with it. Ephemeral identities that never expires, you are it, flag it, and do whatever you want to do with it. And non-privileged identity, suddenly doing admin access, actions, flag it. This is how you climb out of the uncanny valley. You stop worshipping the labels and you start governing the behavior. So, let's land this plane. The uncanny valley of identity categorization is that place where our IM diagrams looks almost perfect, almost human in a way, but the underlying controls just don't line up. We keep adding more types, more subtypes, more OU levels, thinking that if we can just describe reality in enough details, security will magically emerge. But identity in a mature enterprise is not a static noun. It's a living verb, a stream of actions, relationships, and risks over time. If our taxonomies don't translate into distinct, enforced behavior into verbs, then all we have done is create taxonomic debt and really expensive governance theater. The way out isn't another round of renaming accounts. It's shifting our mental model interactive, ephemeral privilege. And then, being brutally honest about what changes in the control plane for each of those. So, if you are up for it, here is my challenge to you. Sometime this week, grab your IM or security teams for 60 minutes. No tools. Just a whiteboard or a dog. List your identity categories. Ask which ones actually change behavior. Circle the ones that are just labels. If you walk out of that meeting with even two or three categories to merge or retire, you have already started paying down taxonomic debt. And once you start seeing identity as a verb, you cannot unsee it. You will catch loopholes earlier. You will design better incentives. And your diagram may get a little less pretty, but your security will quietly get a lot stronger. So, thank you for navigating this one with me. Did you notice the word play there? Navigator? Navigating. If this episode started or sparked ideas, ranty feelings or how know that's us moments, send them my way. You can reach out to me while LinkedIn or you can email me. I love hearing real stories from the trenches from an experts like all of you are. I am Rohit, your identity navigator. And remember, don't fall in love with the label. Follow the behavior. That is where real identity lives.

Podcast Summary

Key Points:

  1. Identity should be viewed as a dynamic, behavioral pattern rather than a static record or noun.
  2. Overly complex taxonomies with deep hierarchies often fail to translate into meaningful, enforced control differences.
  3. The "uncanny valley" of identity categorization occurs when labels look sophisticated but result in identical behaviors and risks across categories.
  4. Real governance comes from asking behavioral questions—such as whether an identity is interactive, ephemeral, or privileged—rather than relying on technical labels.
  5. Categories that don’t map to distinct control behaviors (e.g., MFA, rotation, monitoring) represent taxonomic debt and create governance theater.
  6. Misaligned incentives—like allowing service accounts to bypass MFA—create low-friction backdoors and encourage misuse by humans.
  7. A minimal, behavior-driven taxonomy (interactive vs. non-interactive, ephemeral vs. persistent, privileged vs. non-privileged) enables enforceable, practical control recipes.
  8. The future of identity governance lies in continuous monitoring and automated responses to deviations in behavior, not in complex directory structures.

Summary:

The episode challenges the common practice of treating identity as a static, labeled noun in identity and access management (IAM). It introduces the concept of the "uncanny valley" — where highly detailed taxonomies look mature but fail to translate into actual differences in security controls. The core argument is that identity behaves like a verb, not a record: it's defined by its actions, context, and risk over time.

Instead of relying on complex labels like RPA bots or workload identities, organizations should focus on three behavioral dimensions: interactivity (human vs. machine), lifespan (ephemeral vs. persistent), and privilege (sensitive access vs.

standard). Each combination creates a distinct, governable behavior pattern. The episode emphasizes that taxonomic depth without corresponding control enforcement results in "taxonomic debt"—a gap between label complexity and real security.

A healthier approach is to design only as many categories as can be backed by enforceable policies, then continuously refine based on telemetry and behavior. The ultimate goal is to move from label-based governance to behavior-based automation: detecting anomalies like a non-interactive identity logging in interactively or a temporary identity lasting too long. This shift enables real risk reduction, improves visibility, and prevents security theater.

The challenge to practitioners is to audit their current categories and identify which ones lack real behavioral differences—merging or retiring those that are merely cosmetic. By anchoring identity governance on behavior, organizations create resilient, responsive, and operationally sound systems, avoiding the illusion of maturity that comes from fancy diagrams.

FAQs

It refers to identity taxonomies that look sophisticated and mature but fail to translate into real, enforceable differences in controls or behaviors, creating a gap between appearance and actual security.

It treats identity as a static record in a directory, ignoring that in reality, identity is a dynamic, behavioral pattern over time involving risk, context, and actions.

Is it interactive? Is it ephemeral? Is it privileged? These questions focus on actual behavior rather than static categories like 'service account' or 'RPA bot'.

Interactive identities require MFA, session controls, and phishing resistance, while non-interactive ones use short-lived tokens, strict rotation, and behavioral monitoring.

It signals a policy misconfiguration or security incident, as it represents a breach of behavior expectations and may indicate a backdoor or compromised control.

Taxonomic debt occurs when identity categories are created without corresponding control differences, leading to complexity, unenforced policies, and increased risk over time.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.