Go back

The 10 AI System Design Questions You'll Actually Get Asked in 2026

0m 0s

The 10 AI System Design Questions You'll Actually Get Asked in 2026

This deep dive examines ten AI system design problems asked in 2026 interview loops, framed through four architectural layers: serving, retrieval and memory, trust, and control. The central argument is that classic web architecture fails because it treats compute as elastic, while AI systems are constrained by scarce resources such as GPU memory bandwidth, token budgets, context windows, and human review capacity. The recommended approach is a 90-second triage process asking what is scarce, what failure costs, and what the freshness requirement is, followed by the core rule: name the scarce resource, then design around it. In the serving layer, candidates must explain continuous batching, disaggregated prefill and decode, and two-phase commit quota systems. In retrieval, they must justify structure-aware chunking, hybrid search, cross-encoder re-ranking, and categorized memory with invalidation rules. The trust layer requires layered evaluation, LLM-as-judge mitigations, and decoupled policy from model. The control layer demands durable agent workflows, idempotency, loop detection, and privacy-aware observability. The episode closes with six anti-patterns that signal shallow understanding, concluding that the triage method will outlast any specific technology or question.

Transcription

10864 Words, 62860 Characters

English
0:00 Why Classic Web Design Fails AI Systems I want to start today by putting you, the listener, right in the middle of a room. Like imagine a candidate walking into a system design loop at one of the major AI infrastructure companies. 0:11 Speaker 2 Oh man, the pressure cooker. 0:13 Speaker 1 Yeah, exactly. So we are talking about a senior back end engineer, eight years of experience, deep distributed systems background, like a really strong resume. 0:24 Speaker 2 Right. So they've done the prep. They know the classic cannon coal. 0:26 Speaker 1 Totally. They can design a global news feed or I don't know, on a whiteboard A distributed URL shortener from pure memory with their eyes closed. 0:34 Speaker 2 Yeah, the stuff we all practice for years. 0:36 Speaker 1 Right. So they sit down, grab the dry erase marker and the interviewer gives them the prompt and it goes like this. We serve a large language model to about 2 million daily users. Our GPU bill is roughly $80,000 a day. Design the serving layer. 0:51 Speaker 2 And right there, I mean right there, the trap is set. 0:54 Speaker 1 And this candidate, like so many others, honestly just steps right into it. They go to the whiteboard and they draw a beautifully clean textbook web architecture, like they put a load balancer at the front. They draw application server scaling out a message queue in the middle, a Redis cache sitting in front of a post Chris cluster for the user data, and for scaling they explicitly mention horizontal auto scaling tied to CPU utilization. 1:20 Speaker 2 I mean, look, it is a beautiful architecture. It is a. It's completely logically sound. 1:25 Speaker 1 Right, it would pass a standard web systems around with flying colours. 1:28 Speaker 2 Oh, absolutely. But for this specific interview, it scored an immediate no higher. 1:32 Speaker 1 Yeah, and it's brutal, but understanding exactly why it's a no higher. That is the mission of our deep dive today. 1:39 Speaker 2 Right, because we're two engineers who spend entirely too much time in the trenches of AI infrastructure, and today we're breaking down the 10 AI system design problems that are actively being asked in 2026 loops. 1:52 The Core Principle: Design Around Scarce Resources Yeah, and to understand that poor candidates failure you have to look at the foundational assumption they made. Like every single component they drew treats compute as elastic and cheap. 2:01 Speaker 2 Exactly. In a classic web system, if you get a traffic spike, what do you do? 2:05 Speaker 1 You just spin up another server, you scale out. 2:07 Speaker 2 Right, because that assumption is baked into the DNA of every distributed systems pattern we've learned over the last 20 years. Comute is cheap, network is fast, just sin U more nodes. Yeah, but a GBU cluster is not an elastic web server farm. You cannot just autoscale a Gu cluster in 40 seconds when traffic spikes. 2:27 Speaker 1 No, not at. 2:28 Speaker 2 All you're waiting minutes, sometimes 10s of minutes for capacity that frankly, depending on your cloud region, might not even physically exist at any price. 2:35 Speaker 1 And the detail about auto scaling on CPU utilization, I mean that was the real kill shot. 2:40 Speaker 2 Wasn't it? Oh, that was the nail in the coffin, yeah. You cannot scale an LLM serving layer on CPU utilization because the CPU isn't the bottleneck, right? The thing that saturates, the thing that actually limits your throughput is high bandwidth memory. It is the memory bandwidth on the GPU, not the processor itself. 2:57 Speaker 1 Wow. 2:58 Speaker 2 Yeah, and we can't treat these requests as stateless REST API calls either, right? There's a KV catch a key value cache holding megabytes of attention state in active memory for every single ongoing conversation. 3:10 Speaker 1 So the candidate basically designed around the wrong scarce resource. 3:14 Speaker 2 Exactly. 3:15 Speaker 1 They optimized for request throughput, which is, you know, Web Systems 101. But the interviewer was sitting there with a rubric heavily weighting GPU utilization. That $80,000 a day wasn't just flavor text to set the mood. 3:29 Speaker 2 No, it was the literal constraint of the problem, and the part that should terrify anyone listening who is prepping for these Luke's is that the candidate almost certainly knew all of this. 3:40 Speaker 1 Yeah, you're. 3:41 Speaker 2 Right, if you stopped them in the hallway beforehand and asked hey how does GPU memory bandwidth work or what's AKV cache, they would have given you the right answer. The raw engineering knowledge was absolutely there. 3:51 Speaker 1 But the triage wasn't under pressure, they just pattern matched to the nearest rehearsed shape in their brain, which was a classic elastic web system. They bolted a massive AI model onto a standard web architecture when the interviewer really wanted to see a resource system with a product wrapped around it. 4:07 Speaker 2 Exactly. The candidates failing these staff level rounds aren't failing on intelligence, they are failing on the fundamental framing of the architecture. 4:15 Speaker 1 Which brings us to the core mental model for this entire deep dive. Yes, we are going to look at 4 architectural layers today, the serving layer, the retrieval and memory layer, the trust layer, and the control layer. But before we get to any of them, there is a core rule like a spine to all of this. 4:35 I want you to say it deliberately, because if our listeners take one thing away today, it needs to be this. 4:41 Speaker 2 Name the scarce resource, then design around it. 4:44 Three Questions to Frame Any AI Design Prompt Let's contextualize that for the systems engineers listening in classic system design, the scarce resource was basically always one of three things, right? 4:52 Speaker 2 Yeah, pretty much always. 4:54 Speaker 1 It was disk seeks, network round trips or database connections, caching, sharding, CDNS, read replicas. All those classic techniques are just ways to conserve one of those three resources. This techniques became so standard, we just stopped thinking about the underlying scarcity. 5:09 Speaker 2 Because it never changed. Yeah, but in AI systems the scarcity changes entirely depending on the question, right? Sometimes it's the GPU compute, sometimes it's the token budget, sometimes it's the context window of the model or even human review capacity. If you don't explicitly name which one of those scarcities you were up against in the 1st 2 minutes, you are going to design a beautiful, fully functional, completely wrong system. 5:33 Speaker 1 So how do you stop yourself from falling into that web system trap? 5:36 Speaker 2 You run a specific triage process. This something you as the candidate should do in the 1st 90 seconds of receiving the prompt. Don't touch the whiteboard, just ask yourself and your interviewer 3 questions. 5:49 Speaker 1 OK, the first one is obvious based on what we just said. What is scarce like is it compute tokens context window? 5:58 Speaker 2 Yes, and say it out loud. Literally tell the interviewer the scarce resource here is GPU memory bandwidth, so my architecture is going to center on keeping the GPU busy. 6:08 Speaker 1 Oh wow. Instantly the interviewer knows you've bypassed the trap. 6:11 Speaker 2 Precisely, you've set the stage. 6:13 Speaker 1 OK, so you frame the constraint. What's the second triage question? 6:16 Speaker 2 What does a failure cost? And honestly, this is the one that separates the mid level engineers from the staff and principal levels. 6:23 Speaker 1 Wait, I'm going to push back on that. 6:24 Speaker 2 OK, go for it. 6:25 Speaker 1 Why can't we just handle failures the way we always do? Like, look, in a web system, a failure is a 500 Internal Server Error. You log it, you implement some jittered exponential back off, and you retry, right? The cost of a failure is just one automated network retry. 6:41 Why are we making it more complicated than it needs to be? 6:44 Speaker 2 Because that is a fundamental misunderstanding of what AI systems actually output. You have to look at the wildly asymmetric costs of failure here. 6:53 Speaker 1 Asymmetric how? 6:54 Speaker 2 Well, if a standard REST API fails to load a user profile, sure you retry. But what happens when an LLM fails? It doesn't throw a 500 error, it confidently outputs a hallucination. It succeeds at the HTTP level, but fails at the reality level. 7:10 Speaker 1 Oh right, so it generates text perfectly, but the text is factually bankrupt. 7:14 Speaker 2 Exactly, and the cost of that wrong text depends entirely on the product context you are building for. Let's say you're designing an AI coding assistant. A bad answer there. It's annoying. 7:24 Speaker 1 Yeah, the developer copies it, gets a syntax error in their IDE, rolls their eyes and asks the model to fix it. 7:30 Speaker 2 Right, the cost of failure is low, so you can optimize for speed. But now imagine the prompt is to design a medical summarization tool for patient records. 7:39 Speaker 1 Oh wow, yeah. 7:40 Speaker 2 A false negative, they're say. The model summarizing a chart and quietly omitting a severe penicillin allergy that has a catastrophic, unretrievable real world cost. 7:52 Speaker 1 Right. You can't just exponential back off a medical error exactly. So you literally can't even begin to design your safety pipelines, your fallbacks, or your latency budgets until you know the real world cost of the model being wrong. 8:04 Speaker 2 Precisely, you have to state your assumption about the asymmetric cost of failure out loud. 8:09 Speaker 1 OK. So we have what's scarce and what does a failure cost? What's the third triage question? 8:15 Speaker 2 What is the freshness requirement like? How stale is this model's knowledge allowed to be? Does it need to know things that happened 5 seconds ago? A day ago 1/4. 8:23 Speaker 1 Ago. And why is that a day 0 architectural question? 8:26 Speaker 2 Because it instantly dictates your retrieval strategy and eliminates the most common anti pattern candidates throw out. When asked how to make a model know about recent data, candidates love to confidently say, oh, we'll just fine tune the model. 8:41 Speaker 1 Oh yeah, I've heard that so many times. 8:43 Speaker 2 But fine tuning has a freshness floor measured in days or weeks, right? Think about the mechanics. You have to extract the new data, clean it formatted into instruction pairs, spin up a massive training cluster, run the job, run the offline evils, check for catastrophic forgetting, and then finally deploy the new weights. 9:02 Speaker 1 Which takes forever. 9:03 Speaker 2 Exactly. If the product requirement is that the system needs to know about a user action that happened 5 minutes ago, fine tuning is physically the wrong answer. Knowing the freshness requirement tells you instantly whether you need AR, RAG, pipeline retrieval, augmented generation or if you you can just rely on weights. 9:20 Speaker 1 That is an incredibly powerful framing. 3 questions. What's scarce? What is failure cost? What's the freshness requirement? If you take those 90 seconds, the architecture practically starts designing itself. 9:31 Speaker 2 It really does. 9:32 Optimizing GPU Utilization with Continuous Batching Let's see that in action and dive into our first layer, the serving layer. 9:36 Speaker 2 This is where the money literally burns. 9:38 Speaker 1 Right, when we look at the first set of problems you're going to face, the overarching scarce resource here is GPU memory bandwidth and the actual financial budget. This is that $80,000 a day constraint. So let's put you, the listener, in the hot seat, the interviewer says. 9:55 Design A batched inference API for a GPU cluster. 9:59 Speaker 2 I would argue this is the single most common AI systems question right now. 10:02 Speaker 1 OK, so if I'm pattern matching to my web experience, I hear API and cluster, my brain immediately goes to a standard worker pool pattern. Like a request comes in, hits a queue, a worker node picks it up, claims exclusive access to a GPU, runs the inference, returns the response, and then grabs the next one horizontally scaled. 10:21 Speaker 2 And there goes the interview. Oh no, I mean it sounds entirely logical, but if you draw an exclusive worker pool for LLM inference, you've just proved you don't understand GPU utilization because a single sequence of text being generated cannot come close to saturating the compute capacity of a modern GPU like an H100. 10:41 If you run one request per GPU, you're leaving an enormous amount of highly expensive silicon completely idle. You're paying for a massive cargo ship and using it to deliver a single pizza. 10:53 Speaker 1 That's a great analogy. OK, So what is the staff level move? How do we actually saturate it? 10:59 Speaker 2 You introduce continuous batching. You have to explain to the interviewer that to get efficiency you must bash multiple requests together. But here's the nuance. Traditional batching is like a bus. 11:09 Speaker 1 OK, how so? 11:10 Speaker 2 You wait at the bus stop until 16 passengers get on the bus drives to the end of the line, drops everyone off and comes back. 11:16 Speaker 1 Right, which creates terrible latency in AI because if I asked the model for a quick three word translation and the guy next to me asks for a 2000 word essay, my request is trapped on the bus until his essay is completely finished generating. 11:29 Speaker 2 Exactly. Continuous batching solves this by acting more like an escalator or a rideshare. 11:34 Speaker 1 Oh nice. 11:35 Speaker 2 Request join and leave the active batch dynamically mid flight at token boundaries. When you're short translation finishes on token 5, it drops out of the active batch and a new request from the queue slot to do its exact place for token 6's forward pass. 11:50 Speaker 1 That's brilliant. 11:51 Speaker 2 Yeah, so you explicitly tell the interviewer this is a scheduler problem, not a simple queue problem. 11:57 Speaker 1 That's a great distinction, but to really nail this you have to do the technical deep dive on why we are scheduling it this way. You have to explain the two phases of inference, prefill and decode. 12:07 Speaker 2 Yes, disaggregated prefill and decode is the mark of a senior architect. Let's break the mechanics down. Sure, when a request arrives, the model first has to process the prompt you provided. That is the prefill phase. Think of it like reading a book. It requires massive matrix multiplications, meaning it uses the raw processing power, the FLOPS of the GPU, and it processes all the tokens in the prompt in parallel. 12:31 Got it. It is incredibly compute bounded bursty. 12:34 Speaker 1 But once it digests the prompt, it has to start generating the answer one word at a time. That's the decode phase. Think of that like writing a book. 12:40 Speaker 2 Exactly. And decode has the exact opposite hardware profile. Generating tokens 1 by 1 is completely memory bandwidth bound, right? To generate a single new token, the GPU has to pull the entire massive model, all the weights from the high bandwidth memory, into the compute registers, do a relatively small amount of math and write it back, then do it again for the next token. 13:02 So it is starving for memory bandwidth, not compute. 13:06 Speaker 1 So if you just naively mix prefill tasks and decode tasks on the exact same GPU. 13:11 Speaker 2 They fight each other. A bursty compute heavy prefill task will stall out a steady memory bandwidth heavy decode task. The staff level architect texture separates them. You draw a dedicated prefilled worker nodes that crunch the prompts and then they transfer the KV cash state over the ultra fast network to dedicated decode worker nodes that stream the tokens back to the user. 13:33 Speaker 1 OK, but in system design the real value is knowing the decision boundaries. Like when do you abandon continuous batching? 13:39 Speaker 2 You abandon it when you have hard per customer latency SLA. Say a major financial institution is paying you a premium and your contract guarantees a time to 1st token of under 200 milliseconds unconditioned. 13:51 Speaker 1 Right. And you can't guarantee that in a shared continuous batching pool, because if the scheduler is busy packing a massive batch for free tier users, the bank's request might sit in the queue for a few milliseconds, or the prefill might take slightly longer due to batch size. 14:06 Speaker 2 Right. In that scenario, you look the interviewer in the eye and articulate the trade off. For our enterprise tier, we will provision dedicated unbatched or strictly isolated GPU capacity. Yeah, our hardware utilization will drop terribly, Our cost per request will skyrocket, but we are actively accepting that financial trade off to meet the strict latency SLA and. 14:29 Speaker 1 When you're talking about capacity, you need to throw out specific numbers to show you know the metal, like the KV cash footprint. 14:35 Speaker 2 Yes, exactly. The interviewer wants to know if you know what actually causes an out of memory error on the GPU. Junior candidates think it's the model weights. No, the weights are static. A 70 billion parameter model takes up the same amount of VRAM whether it's serving 1 user or 50. 14:52 What actually scales and crashes the node is the KV cache. You explained that the KV cache memory footprint scales as a function of the batch size multiplied by the sequence length. If you let users send massive 100K token prompts and you try to batch too many of them, the KV cache will consume all the VRAM and crash the node long before you hit any compute limits. 15:12 Speaker 1 All right, so we figured out how to optimize the compute, but how do we charge for it? 15:17 Metering Unknown Consumption with Two-Phase Commit The interviewer pivots and says design A token metered rate limiter. 15:21 Speaker 2 Again, we're in the serving layer. The scarce resource is still compute and dollars. 15:26 Speaker 1 No, a rate limiter is classic distributed systems. I built these. I'd naturally suggest a standard token bucket algorithm or a sliding window request per second model backed by Redis. 15:36 Speaker 2 Yeah, that's what most people say. 15:37 Speaker 1 You give the user an API key. They get 100 requests a minute. If they hit one O 1, return a 429 status code. Done. 15:44 Speaker 2 And if you say that the interviewer writes down that you missed the fundamental paradigm shift of AIAPIS. You are proposing metering a resource whose consumption is entirely unknown at admission time. 15:57 Speaker 1 Walk me through that. Why is it unknown? 15:59 Speaker 2 Think about a traditional API endpoint. If you hit APL's profile, the database query costs roughly the same amount of CPU and memory every single time, so metering by requests per second works perfectly. One request equals 1 unit of work. 16:14 Speaker 1 OK, Yeah. 16:15 Speaker 2 But in an LLMAPIA user hits your generate endpoint, one request might have a short prompt asking for A1 sentence summary, consuming maybe 12 tokens, right? The exact same endpoint might get hit a second later by a request uploading A dense legal contract asking for a comprehensive analysis consuming 4000 tokens. 16:33 Speaker 1 300 times the compute cost, but it registers as just one request to a standard Redis sliding window. 16:39 Speaker 2 Exactly. If you just use RPSA, heavy user can completely bankrupt your compute cluster while staying perfectly within the request limits. You cannot measure the true token consumption until after the request is finished running. 16:51 Speaker 1 So what's the architectural move? How do you rate limit based on a number you literally don't know yet? 16:56 Speaker 2 You introduce A2 phase commit against a central quota. OK. When the request arrives, you perform admission control based on a maximum estimated cost. You look at the prompt length and you look at the Max tokens parameter the user is forced to pass in. You go to your quota service and say reserve a budget for 4000 tokens. 17:15 If they have the balance, you hold those tokens and admit the request. 17:18 Speaker 1 And when the request finishes. 17:20 Speaker 2 Let's say it only generated 500 tokens before hitting a stop sequence. You trigger the second phase, you go back to the quota service and release the unused 3500 tokens back to their available balance. 17:31 Speaker 1 OK, that makes sense for a standard request response loop. But wait, what about streaming? Most AI products stream the tokens back to the user in real time so they aren't staring at a blank screen for 30 seconds, right? If you're streaming, you are spending that token budget one word at a time mid flight. 17:49 Speaker 2 That is the exact complication the interviewer will throw at you. Streaming requires mid flight stream cutting your rate limiter. Can't just live at the API gateway anymore. Checking things at the start and end. 18:01 Speaker 1 So where does it live? 18:02 Speaker 2 It has to be integrated closely with the inference engine itself. You need a background processor or a sidecar that eriodically checks the accumulated token count of the active stream against the remaining quota. If they run out of budget midsentence, the system has to gracefully sever the connection. 18:19 Speaker 1 And you need a specific user facing error policy for that. You can't just drop the TCP connection because the client app will just think the Internet flipped and try to auto retry. You have to send a final Jason payload that explicitly says stream terminated. 18:33 Speaker 2 Quota exceeded precisely. Now, what's the decision boundary for this? Do you use a centralized quota service like a global Redis cluster or do you do local per node limiting? 18:44 Speaker 1 Well, if it's a centralized service, you're introducing a network round trip for every single admission check. That adds latency to the time to 1st token. 18:52 Speaker 2 Yes. So the alternative is local limiting. You give each API node its own isolated slice of the user's budget. But what happens if a user's traffic is distributed across 20 different nodes via the load balancer? 19:06 Speaker 1 Oh, I see. 19:07 Speaker 2 They could theoretically burst and drain all 20 local buckets simultaneously, heavily overshooting their actual financial quota before the nodes sync up. 19:16 Speaker 1 I see. So the boundary is financial risk tolerance. When independent local buckets allow too much financial overshoot across nodes, When the cost of that leak compute is unacceptable to your finance team, you abandoned local buckets. You accept the latency hit of a network round trip to a centralized quota service to ensure strict financial enforcement. 19:35 Speaker 2 You articulate that trade off, you pass the question. 19:38 Speaker 1 All right, third scenario in the serving layer, we've got a compute scheduled, we've got a rate limits now. 19:42 Dynamic Model Routing and Semantic Caching Risks The interviewer asks design A multi tenant model gateway with cost routing. 19:47 Speaker 2 This tests real world product economics. You don't just have one model in production. You have a frontier model, the massive, brilliant, expensive 100 billion parameter 1. You have a mid tier model and you have a small lightning fast cheap 8 billion parameter model. 20:04 Traffic is coming in from thousands of users with vastly different needs. 20:08 Speaker 1 A trap here seems pretty obvious. It's static config routing. Endpoint A maps to the frontier model, Endpoint B maps to the small model. The client A just decides where to send the request. 20:19 Speaker 2 Easy. Easy to build, but it leaves massive savings on the table. A user might send a trivial prompt like What is the capital of France? To the frontier model, costing you a fortune in compute for something that tiny model could answer perfectly for fractions of ascent. 20:33 Speaker 1 So the move is to treat routing as a dynamic policy layer. You build a gateway that intercepts the prompt and classifieds the difficulty of the request cheaply. 20:42 Speaker 2 Exactly. But how cheaply you start with heuristics. What's the raw length of the prompt? Does it contain complex markdown or code snippets? Is it a known complex task type? 20:50 Speaker 1 And if heuristics aren't enough. 20:52 Speaker 2 If heuristics aren't enough, then you might route it through a tiny sub billion parameter classifier model. The goal is to send the request to the cheapest model predicted to succeed. 21:02 Speaker 1 But wait, how do you know if it succeeded? If you route a complex coding question to the small model and it gives a terrible hallucinated answer, the user just gets a terrible answer. You saved money, but you ruined the user experience. 21:16 Speaker 2 And that is the crux of the staff level answer. The router is useless without a verification signal. You have to define what succeed means and how you detect failure fast enough to escalate. Maybe you write it to the small model, but you have a lightweight fall back trigger. 21:32 If the small model outputs an uncertainty token, or if you get stuck in a repetitive loop, you immediately abort and transparently reroute the request to the frontier model. A dynamic router without a verification signal is just a coin flip of the budget. 21:45 Speaker 1 OK, I have an idea for saving money here. What if we just aggressively semantic cache everything? 21:50 Speaker 2 Hear me out. If a prompt comes in, we embed it, do a similarity search in a vector database against AST prompts, and if it's 95% similar we just return the cache answer that bypasses the models entirely. Massive savings. 22:04 Speaker 1 If you suggest that, I will push back on you extremely hard and so will your interviewer. 22:09 Speaker 2 Why everyone talks about semantic caching? It's the hottest buzzword right now. 22:13 Speaker 1 Match caching is fine if the prompt string is identical byte for byte cache away. But semantic caching on embedding similarity is highly dangerous in a gateway. Think about the mechanics of embeddings. Let's say you have a medical lab. Prompt A is patient has a history of high blood pressure, what is the recommended dosage for drug X? 22:33 Prompt B is patient has no history of high blood pressure. What is the recommended dosage for drug X? 22:38 Speaker 2 Oh wow. In vector space those two sentences are going to be incredibly close. The embeddings are nearly identical because the vocabulary is identical. The cosine similarity is going to be like .98. Yes, they are near duplicates mathematically, but they require fundamentally different, potentially life altering answers. 22:56 If your semantic cache threshold is too forgiving, you will confidently serve the wrong dosage. To the second patient, you risk serving confidently wrong cache data. You must name that risk before the interviewer does. 23:07 Speaker 1 So what's the decision boundary for dynamic cost routing overall? When do you just go back to the static config? 23:14 Speaker 2 You drop dynamic routing when your traffic is highly homogeneous. If you are building a specialized tool that only does deep complex code refactoring, then 99% of requests legitimately need the Frontier model anyway. Yeah, in that case the dynamic router is just adding latency, complexity and its own failure modes without actually saving you any money. 23:33 You drop the complex architecture when the product reality doesn't justify it. 23:36 Speaker 1 That makes total sense. 23:37 Designing RAG for Scale: Chunking, Hybrid, Re-ranking OK, let's take a breath. We've covered the serving layer, and I want to pause to reflect on the pattern here. In all three of those scenarios, the trap was relying on an elastic compute web pattern, and to avoid it, we go back to the spine of our deep dive. 23:51 Speaker 2 Name the scarce resource, then design around it. 23:55 Speaker 1 In those first three questions, the scarce resource was GPU bandwidth and money. But as we transition to the retrieval and memory layer, the scarcity completely shifts. We aren't worried about the GPU crunching the tokens anymore. 24:07 Speaker 2 Right. In this layer, the scarce resource's attention, specifically the context window of the model and what earns the right to take up space inside it. The context window is strictly finite. Yeah. Even if a model has a massive million token context window, packing it full of garbage degrades the model's ability to reason, dilutes the focus, and costs a fortune in compute. 24:31 Speaker 1 Every architectural decision in this layer is about curation, defending the context window. So the interviewer says design A rag pipeline over 50 million documents. 24:40 Speaker 2 RAG retrieval Augmented generation. 24:43 Speaker 1 The tutorial standard is always the same, and it's definitely the trap here. You take all 50 million documents. You split them into 512 token chunks. You run them through an embedding model. You dump those dense vectors into a vector database. 24:54 Speaker 2 And so easy. 24:54 Speaker 1 Right, when a user asks a question, you embed the question, do a top cosine similarity search in the database, grab the top five chunks, stuff them into the prompt, and hit generate. 25:05 Speaker 2 And if you describe that pipeline in a staff interview, you have just given a junior level answer. It's the tutorial architecture and it falls apart immediately at scale. The staff level reality requires breaking down for specific design decisions first. Chunking. 25:21 Speaker 1 Right, splitting it by 512 tokens. 25:24 Speaker 2 Which is terrible. Chunking is a massive design decision. If you use fixed arbitrary token counts, you might split a document right in the middle of a crucial paragraph, or worse, in the middle of a Python function. When the retriever grabs the second-half of that split, it lacks all the preceding context. 25:41 The model gets a chunk of code with no variable declarations. You tell the interviewer that you will chunk by structure. 25:47 Speaker 1 OK structure, how you? 25:48 Speaker 2 Use an AST parser, an abstract syntax tree parser for code so it knows to keep a def block together. Or you use a markdown parser for text to ensure that headers, paragraphs and logical blocks stay intact as a single semantic unit. 26:01 Speaker 1 OK, structure over arbitrary tokens. What's the second design decision? 26:07 Speaker 2 Hybrid retrieval relying purely on dense vector similarity is a massive tell that you haven't put one of these in production. 26:14 Speaker 1 Really. Why? 26:15 Speaker 2 Vectors are great for conceptual matching and paraphrasing, but they are notoriously bad at exact keyword matches. If a user searches for a specific serial number like error code 4X99BA vector search might just return general conceptual documents about system errors because the specific string gets lost in the dense embedding AH. 26:37 Speaker 1 I see. 26:37 Speaker 2 You have to explicitly state that you are mandating hybrid retrieval. You run the dense vector search alongside a traditional sparse BM25 keyword search, and you combine the candidate pools. 26:48 Speaker 1 And that leads into the Third Point which is re ranking. 26:51 Speaker 2 Exactly because retrievers, both vector and keyword, are optimized for recall. They are fast and they are good at finding a large pool of potentially relevant chunks from ACE of 50 million documents. But they aren't smart enough to know which ones are the most relevant, right? Think of the retriever like an intern. 27:07 You send them into a massive library to find books about a topic. They come back with 100 books that seem relevant. They did their job fast. 27:14 Speaker 1 But you can't shove 100 books into the context window. 27:17 Speaker 2 Right, so you pass those hundred documents to a re ranker, usually a cross encoder model. The reranker is the senior researcher. It is optimized for precision. Cross encoder doesn't just look at 2 separate embeddings, it concatenates the user's query and the document chunk together and runs full self attention across both of them simultaneously. 27:38 It does a deep, computationally expensive comparison to understand exactly how the query interacts with the text and picks the absolute best 5 chunks to actually put in the context window. 27:48 Speaker 1 But here is a crucial operational question. You've built this complex pipeline, AST chunking, hybrid retrieval, cross encoder, re ranking, and then LLM generation. If the user gets a terrible answer, how do you even debug that? Was the LLM generation bad or did the retriever just fetch the wrong documents? 28:05 Speaker 2 That is exactly what the interviewer is waiting to hear you address. You cannot rely on Vibes to know if retrieval is working. You have to introduce retrieval specific metrics that are completely decoupled from the LLMS generation quality. 28:18 Speaker 1 Like what? What's the math? 28:20 Speaker 2 Metrics like recall at K and mean reciprocal rank or Mr. recall at K measures if the highly relevant golden document was actually present and the top key results returned by the intern. MRR measures how high up in the ranked list the correct document appeared. 28:36 OK, got it. If the correct document is consistently ranked 50th, your MRR is terrible and your cross encoder isn't doing its job. You have to build an evaluation set of queries and known good documents and constantly measure your retriever against it offline. 28:50 Speaker 1 So what's the decision boundary here? When do you stop using R entirely? 28:54 Speaker 2 You shift away from rag and move to fine tuning when the product requirement shifts from facts to behavior or format. 29:01 Speaker 1 Ari is incredible for injecting specific up to 29:23 You articulate that boundary clearly. That's a great distinction. 29:27 Four Architectural Buckets for Long-Term Memory Let's move to the next scenario in the memory layer. Design chat memory across sessions. 29:33 Speaker 2 This is deceptively hard. Candidates vastly underestimate how complex long term memory is. 29:38 Speaker 1 The trap is obvious here too. The user wants the AI to remember them, so you just take the entire conversation history, store it in postgres, and every time they send a new message, you pull the whole history, slap it into the prompt, and hit generate. 29:49 Speaker 2 And it works beautifully at Turn 3 of the conversation. But what happens at Turn 906 months later? That is where real products live. Yeah, if you try to the stuff 900 turns into the prompt, you will blow past the context window limit. Your latency will degrade to dozens of seconds and your inference costs will bankrupt the company you are wasting the scarce resource context space on. 30:10 Hi, how are you from six months ago? 30:12 Speaker 1 So what is the staff level move? 30:14 Speaker 2 You categorize memory into 4 distinct architectural buckets before you even think about designing the storage layer. First you have the working set. This is the immediate context. The last five or ten turns of the conversation. It needs to be verbatim fast and always in the prompt. 30:29 Speaker 1 Makes sense. What's the second? 30:30 Speaker 2 Durable facts. These are core pieces of information the model learns about the user over time. The user's name is Sarah. They prefer Python over Java. They have a peanut allergy. 30:42 Speaker 1 Right, things you always need to know. 30:44 Speaker 2 Yeah, these should not be left in the raw transcript. You run an asynchronous background process, maybe a smaller LLM, that reads the recent transcript, extracts these facts and stores them structurally in a database. Then you inject only the relevant ones into the system prompt. 30:59 Speaker 1 OK, working set and durable facts third. 31:02 Speaker 2 Episodic recall. This is basically a localized our job system over the user's personal history. If the user asks what was that marketing strategy we brainstormed last March, the system embedding searches the historical conversation logs and retrieves just those specific chunks to answer the question. 31:19 Speaker 1 Very cool. And the 4th. 31:21 Speaker 2 And finally, the 4th bucket is summary state. This is a highly compressed rolling summary of everything too old to keep in the working set, but too broad to be a single durable fact. 31:32 Speaker 1 Wow, 4 totally different mechanisms. But wait, I have a challenge for the durable facts bucket. Let's say the background model extracts A durable fact. The user lives in Chicago, it stores it, but then a year later the user says they're moving to Seattle. 31:48 If the system is still injecting user lives in Chicago into the system prompt, a stale fact is actually worse than having no memory at all, because the model will hallucinate confidently based on bad data. 31:59 Speaker 2 That is the exact edge case you need to bring up. You must design invalidation and precedence rules. You have to architect a system where a recent extraction can explicitly overwrite or contradict an old extraction. 32:11 Speaker 1 How do you do that? 32:12 Speaker 2 You usually do this by running a semantic search on the existing facts before storing a new one and prompting the extraction model to resolve conflicts. And beyond that you need to propose user editable memory. 32:22 Speaker 1 Like AUI? 32:23 Speaker 2 Exactly. You build AUI where the user can see what the AI knows about them and manually delete or edit those facts if you propose that you have just answered a deeply technical systems question and a complex product question in one move. 32:38 Speaker 1 OK, let's hit the last scenario. 32:39 Fusing Hybrid Retrieval Scores with RRF In the retrieval layer, the interviewer wants to dig into the math design. Hybrid retrieval with re ranking at scale. 32:48 Speaker 2 We touched on this earlier, combining keyword search like BM25 and vector search like cosine similarity, but the trap here is when the interviewer asks how exactly you combine them. The naive answer is 0. We'll just run both searches, get the scores, and average the results. 33:03 Speaker 1 Which sounds perfectly. 33:04 Speaker 2 Reasonable. It sounds reasonable, but it is mathematically flawed. It shows a lack of deep understanding of how these algorithms actually score documents. Why? Because you cannot average scores from entirely different mathematical distributions. ABM 25 keyword score is completely unbounded. Depending on term frequency, it might return a score of 150 for a really good match, or 12 for a mediocre 1. 33:25 Speaker 1 OK and vector score. 33:26 Speaker 2 A cosine similarity vector score is strictly bounded, usually between zero and one. If you try to average a score of 150 with a score of .85, the keyword search will completely dominate the vector search every single time. Averaging them is meaningless. 33:42 Speaker 1 So how do you actually fuse them? 33:44 Speaker 2 You have to name specific normalization techniques. You either explain that you would do per source score normalization, mapping both distributions to a standard scale using min Max scaling before weighing them, or more likely you mentioned reciprocal rank fusion or RRF. 33:59 Speaker 1 Walk me through the math of RRF. How does it work? 34:02 Speaker 2 RRF ignores the raw scores entirely. It only looks at the position of the document in the ranked lists. If document A was ranked one in the vector search and ranked four in the keyword search, RRF calculates a new score based purely on those ranks. Interesting. The formula IS1K plus rank. 34:18 It penalizes documents that only appear in one list, and highly rewards documents that appear near the top of both lists, regardless of the wildly different raw scores the underlying systems generated. 34:28 Speaker 1 That is fascinating. And what's the decision boundary for this entire re ranking process? Because running a cross encoder over 100 documents every time a user searches sounds incredibly expensive and. 34:38 Speaker 2 That is exactly the boundary. You drop the heavy cross encoder re ranker when the latency cost violates your strict P99 latency budget. If you are building a real time auto complete feature, you absolutely cannot afford the 300 milliseconds it takes to run self attention over a large candidate set. 34:57 Speaker 1 Right, you just drop back to simple RRF fusion. 35:00 Speaker 2 Exactly. You have to know roughly what these operations cost in time. Retending that reranking 100 documents is computationally free is a massive red flag. 35:08 Speaker 1 That wraps up the retrieval and memory layer. The scarce resource was context, and every decision was about fiercely rotecting what gets into that finite window. 35:16 Speaker 2 Exactly. 35:17 Testing Nondeterministic Systems with LLM as Judge Which brings us to our third layer, the trust layer. 35:20 Speaker 1 Here the scarcity shifts again. It's no longer compute and it's no longer the context window. The scarce resource here is confidence, knowing that the non deterministic system you build actually works and the finite capacity of human reviewers. So the scenario design and evaluation harness for an LLM feature. 35:39 Speaker 2 I cannot stress this enough, this is the most under prepared question in these interviews. 35:44 Speaker 1 I believe it. It breaks standard software engineering paradigms like we are all taught how to test deterministic systems. You write a unit test, assert, add two 2 = 4. If it equals 4, the test passes. If it equals 5, it fails. It's binary, but LLMS are fluid. 36:00 If you give an LLM the exact same prompt 10 times, you might get 10 slightly different phrasing variations. How do you write a unit test for that? 36:07 Speaker 2 The trap candidates fall into is treating it as an operational QA problem, not a systems engineering problem. They say we'll just log a sample of the outputs in production and we'll have human annotators review them to make sure they are good. 36:19 Speaker 1 Which doesn't scale at all. If I'm a developer and I tweak a system prompt, I can't wait three days for a human QA team to review 5000 outputs before I merge my pull request. 36:29 Speaker 2 Right, it is not an architecture. The staff level evaluation architecture requires you to outline 3 distinct layers. Layer one is your fixed benchmark set. This is your CICD regression suite. You curate a golden data set of highly diverse, challenging inputs alongside known good human verified outputs or rubrics. 36:50 Every time a developer tweaks a prompt or updates a model weight, this suite runs automatically gives you a baseline. Without it, you cannot scientifically tell an improvement from a regression. 37:00 Speaker 1 OK, but how do you grade the new outputs against the golden ones? If the text is slightly different, you can't just do an exact string match. 37:06 Speaker 2 That brings us to layer 2 LLM. As a judge, you use a massive frontier model like a G PT4 or Claude 3 Opus to grade the output of your smaller production model. Really. Yes, you give the judge a strict, well engineered grading rubric in its system, prompt the golden answer and the generated answer, and you ask it to score the generation on a scale of 1 to 5 for accuracy, tone and safety. 37:29 Speaker 1 I have to push back here. Isn't LLM as a judge inherently flawed? We know these models have massive biases. They suffer from verbosity bias, they automatically think a longer rambling answer is a better answer, and they have self preference bias. They will rate answers generated by their own model family higher than competitors. 37:49 Speaker 2 You are absolutely right, and bringing up those specific flaws is exactly how you score points. You don't resent LLM as judge as a magic bullet, you resent it as a flawed tool that requires engineering mitigations. 38:01 Speaker 1 OK, wait. 38:01 Speaker 2 What to fix osition bias? You randomize the order of the answers you resent to the judge to fix selfreference. You never use the exact same model family as both the generator and the sole judge. And most importantly, you continually calibrate the LLM judge's scores against a small rotating sample of human graded outputs to ensure its grading actually aligns with human reality. 38:22 Speaker 1 OK, so layer one is the benchmark suite, Layer 2 is LLM as judge. What's the third? 38:28 Speaker 2 Layer 3 is online signals telemetry from the real world. Thumbs up, thumbs down buttons, copy to clipboard events, user abandonment midstream. 38:37 Speaker 1 Makes sense? 38:38 Speaker 2 Individually, these are noisy signals. A user might click thumbs down just because they didn't like the fact the AI told them, not because the AI was factually wrong, but in aggregate over millions of requests. Online signals are the only metric that actually measures reality. 38:53 Speaker 1 So you have these three layers generating tons of data, but there's a trap with the metrics themselves because looking at an aggregate dashboard is dangerous. 39:01 Speaker 2 Aggregate metrics lie. Let's say your dashboard says your LLM is succeeding 94% of the time. You think, great, chip it, yeah. But if you don't slice that data, you are completely blind. You might be succeeding 99% of the time in English, but only 60% of the time in Spanish. Or you might be doing great on short queries but failing catastrophically on complex reasoning queries. 39:22 A staff engineer explicitly states that they will architect their telemetry to slice the evaluation metrics across multiple cohorts, languages, and query types. 39:31 Speaker 1 Let's move to the other side of the trust layer. Design A content moderation pipeline for a chat product. 39:37 Decoupling Policy from Model in Safety Pipelines The trap here is so common. The candidate goes to the whiteboard and draws a box labeled Safety Classifier right in front of the LLM request comes in, hits the classifier. If it's bad, block it. If it's good, send it to the model. Done. What's? 39:51 Speaker 1 Wrong with that. 39:52 Speaker 2 It completes 2 entirely different things, the mechanism of the policy. The highest signal move you can make on this question is to explicitly separate policy from the model. 40:00 Speaker 1 Break that down. 40:00 Speaker 2 Policy is what is not allowed. It is business rules. It dictates that we do not allow hate speech or self harm instructions or whatever the legal requirements are. The policy is owned by a trust and safety organization, OK. It needs to be versioned, auditable and it changes on the legal and organizational cadence. 40:17 If a new digital safety law passes in the EU today, the policy has to update overnight. 40:22 Speaker 1 Then the model. 40:23 Speaker 2 The model that the classifier is just how you detect it. It is trained on data, and it changes on a slow machine learning cadence. It might take weeks to retrain and evaluate a new safety classifier if you tightly couple the policy logic into the model weights itself. You cannot change a business rule without deploying a massive ML release. 40:41 Oh. 40:42 Speaker 1 That sounds like a nightmare. 40:43 Speaker 2 And worse, you cannot explain to a regulator why a specific decision was made because the reasoning is locked inside a black box neural network. 40:51 Speaker 1 So you decouple them. The classifier just outputs semantic tags or probability scores like this is 80% likely to be toxic and a separate deterministic rules engine. The policy layer actually makes the business decision on whether to block it based on those tags. 41:05 Speaker 2 Exactly. And once you have that, you design A risk tiered cascade. You don't run every single message through an expensive heavy moderation model. You run a cheap, fast, maybe even rejects based classifier and everything. For the vast majority of normal traffic, it passes instantly. 41:22 For the obviously egregious stuff, it blocks instantly. It's only that uncertain Gray area middle band that you escalate to a heavier, more expensive review process, either a larger LLM or a human queue. 41:34 Speaker 1 What about streaming? If the user is watching the text appear word by word on their screen, how do you moderate that? 41:41 Speaker 2 That is a fascinating visibility boundary problem. The architectural question is what content is allowed to become externally visible on the user's Dom before the safety decision has been finalized. Once a toxic token is rendered on the screen, the damage is done. 41:57 You cannot unsay it. You have to buffer the stream slightly running parallel classification on rolling windows, and you have to architect a way to cut the stream and clear the UI if the threshold is breached mid sentence. 42:08 Speaker 1 And there's a human element here that we can't ignore. If you are escalating Gray area content to human reviewers, you are exposing your own employees to potentially horrific. 42:17 Speaker 2 Yes, designing for reviewer welfare is an architectural concern. You must build systems that redact obvious PII automatically, that blur images by default, that strip audio, and that throttle the amount of traumatic content any single reviewer sees in a shift. 42:34 If you mentioned that you are designing the queue to reduce human exposure, interviewers at mature safety forward companies will absolutely notice. 42:42 Speaker 1 What is the decision boundary for moderation? 42:44 Speaker 2 It's the fail open versus fail close dilemma. If your safety classifier service goes down, and it will, do you fail open, allowing all messages through so the product stays up, or do you fail closed, shutting down the entire chat product to ensure no unsafe content leaks? 42:59 Speaker 1 I assume it depends entirely on the product. Like if it's a pediatric mental health chat bot, you fail closed, the risk of harm is too high. If it's an internal coding assistant for senior developers, you fail open to maintain velocity. 43:11 Speaker 2 Exactly. The key is that you articulate the trade off. Suggesting A degraded mode classifier local lightweight filter that kicks in during outages is a better staff level answer than just flipping a coin between open and closed. 43:22 Durable Workflows for Idempotent Agent Actions OK, we are entering the final layer, the control layer. These are the newest hardest questions on the circuit. These are classic distributed systems problems that are wearing an AI costume. 43:33 Speaker 2 And this is where the scarce resource becomes state and observability. 43:36 Speaker 1 Scenario design and agent execution runtime with durable tool calls. Let's set the stage for people who haven't built agents yet. You have an LLM that doesn't just chat, it executes plans. It thinks it decides to call an external API a tool, it looks at the result and it loops. 43:54 It might do this for minutes or even hours across 50 different steps. 43:58 Speaker 2 And the trap is looking at that and saying, oh I'll just write a while loop in a Python script. 44:02 Speaker 1 Which is exactly what every beginner tutorial does. While not finished. Just do step. Why is that a trap? 44:08 Speaker 2 Because it completely ignores the reality of distributed systems. What happens if your server crashes at step 40? 44:14 Speaker 1 Well, if it crashes, wait, why not just restart the script? Just run the loop again from the beginning. 44:18 Speaker 2 Because Step 3 might have been charge the user's credit card via Stripe or drop a database table. If you just restart a blind while loop, you are going to execute destructive non idempotent actions multiple times. 44:33 Speaker 1 Ouch. Yeah, charging a customer twice is a resume generating event. So how do you architect it? 44:38 Speaker 2 You reframe the problem. An agent is not a script, it is a durable workflow. Think Temporal or AWS step functions. Every single step the agent takes, every thought, every tool call, every observation is a state transition in a state machine, and every transition must be checkpointed to a durable database like Postgres. 44:58 If the process dies at step 40, the system must be able to wake up on a completely different server, read the checkpoint, and resume a step 41 without repeating anything. 45:07 Speaker 1 And to handle the credit card problem. 45:09 Speaker 2 You introduce idem potency keys on every single tool call. The agent runtime automatically generates you need hash for that specific step in the plan and passes it in the header to the payment API. That way, even if a network partition causes a retry, the API knows it's the same request and doesn't double charge. 45:25 This is where your classic distributed systems knowledge pays massive dividends. You just have to know to apply it here. 45:31 Speaker 1 What about AI specific issues in these loops? Agents are notoriously dumb. Sometimes they get stuck. 45:37 Speaker 2 Yes, loop detection is critical. Agents will hallucinate a bad API call, get a 400 error, and then just keep trying the exact same bad call infinitely. You have to design the runtime with step budget, say maximum 20 steps per task. 45:53 And you need repetition detection algorithms that force the agent to pause and ask a human for help if it repeats the same action three times. 46:01 Speaker 1 There's also a context roblem here. If the agent runs for 50 steps an every API response is getting appended to the prompt, aren't you going to run out of context window? 46:10 Speaker 2 Absolutely. The context will bloat until the API rejects the request. A staff level architecture includes a context compaction strategy. You need a background process, usually a smaller cheaper model that summarizes older steps of the workflow, keeping only the high level intent and the most recent granular observations. 46:26 Otherwise the agent hits a wall mid task and literally forgets what its original goal was. 46:30 Speaker 1 What's the decision boundary for all this heavy durability? 46:33 Speaker 2 Overhead durability is expensive. Database rights are slow. You drop the durable runtime and go back to a simple in memory loop when the tasks are very short, completely side effect free, and cheap to redo. If the aging is just reading a few internal documents to summarize them, you don't need a heavy temporal workflow. 46:52 Just run it in memory. If it fails, run it again. 46:54 Tracing, Cost, and Privacy in AI Observability All right, the final scenario, design observability for an LLM powered system. 46:59 Speaker 2 This is one that nobody was asking two years ago, but it is mission critical now. 47:03 Speaker 1 The trap is obvious. I'll use betadog. I'll log all the requests. I'll log the responses. I'll put up a dashboard showing P99 latency and 500 error rates. 47:12 Speaker 2 That is necessary, but is nowhere near sufficient. You have to explain to the interviewer the fundamental gap in observability. Here. Traditional observability answers one question. Did it work? Did the server return a 200 OK? Did it return fast? 47:26 Speaker 1 But AI observability has to answer something entirely different. 47:29 Speaker 2 It has to answer was it any good? And LLM can return a beautiful 200 OK status code in 50 milliseconds while confidently hallucinating absolute garbage that insults your user. If you just look at a traditional dashboard, that system looks completely healthy. 47:46 Speaker 1 So how do you actually build observability that catches that? 47:49 Speaker 2 You have to trace the full chain. Yeah, an LLM response isn't just one API call, it's a pipeline. It's the user input, the retrieval step, the re ranking, the prompt construction, the generation, the tool calls, and the post processing. OK, if a user gets a bad answer, you need distributed tracing that connects all those spans. 48:06 Was the answer bad because the LLM hallucinated or because the RH pipeline retrieved the wrong document? If you don't have tracing, you can't attribute the failure to a stage and. 48:15 Speaker 1 What about costs? 48:16 Speaker 2 Cost per request is a first class metric in AI. It belongs right next to latency on your dashboard. Because 1 user might trigger a workflow that costs 1000 times more than another user. You have to admit telemetry on token counts and estimated cost for every single span in your trace. 48:30 Speaker 1 But tracing full prompts and outputs, logging exactly what the user typed and what the model said, That sounds like a massive privacy nightmare. 48:38 Speaker 2 It is, and if you don't bring it up you fail the privacy check. Prompts contain sensitive user data passwords. PII proprietary kite. You cannot just dump full payloads into your logging vendor. You have to design redaction proxies that strip PII before it leaves your network. 48:55 You get strict retention limits and role based access control on the logs themselves. 48:59 Speaker 1 So what's the decision boundary on logging? Do you log everything? 49:03 Speaker 2 You can't. Full capture of every prompt and output for millions of users is a privacy liability and a staggering storage bill. The boundary is your sampling rate. You tie the sampling rate dynamically to your error signals for successful normal interactions. Maybe you only log 1% of full traces, but if the user clicks the thumbs down button or if the model triggers a safety filter, you instantly adjust the sampling rate for that session to 100%. 49:27 You balance debug ability against the storage bill. 49:29 Speaker 1 That's brilliant. OK, 10 questions, 4 layers, the serving layer, retrieval, trust and control. 49:36 Common Mistakes and How to Avoid Them And through all of it, the same spine kept appearing. 49:39 Speaker 2 Name the scarce resource, then design around it. Whether it was GPU memory token budgets, context windows or human review capacity naming, the constraints solve the architecture. 49:52 Speaker 1 Now we are going to rapid fire the six anti patterns. These are the six sentences that you hear in interviews that immediately signal the candidate doesn't understand the domain. 50:01 Speaker 2 I've heard all six in real interviews. Let's run through them. 50:04 Speaker 1 Anti pattern one. We'll just fine tune the model. 50:06 Speaker 2 Almost always said when the candidate is trying to solve a knowledge problem. Fine tuning is for shaping behavior, tone and output format. It is a terrible, slow and expensive instrument for injecting factual knowledge. It instantly fails the freshness question. Say R Reg instead. 50:22 Speaker 1 Anti pattern 2 will cache the responses. 50:24 Speaker 2 The fatal flaw is not specifying which cache. You must distinguish between exact match caching, which is safe but low yield, and semantic caching, which is high yield but carries massive staleness and correctness risks like the blood pressure example. 50:38 Speaker 1 Anti pattern 3 will add a classifier. 50:41 Speaker 2 The content moderation trap. A classifier is a mechanism, not a policy system. Saying add a classifier proves you haven't thought about the policy layer, the escalation thresholds, or the trust and safety decoupling. 50:53 Speaker 1 Anti pattern 4 will use a vector database. 50:56 Speaker 2 This one drives me crazy. It's using a vendor name. When a retrieval strategy belongs there, it means nothing. If you say vector database but say nothing about how you chunk the documents via AST or how you perform hybrid keyword search, you've just named a noun instead of making a design decision. 51:13 Speaker 1 Anti pattern 5 will auto scale the GPU's. 51:16 Speaker 2 The cold open mistake. GPU capacity is not elastic on the time scales of a web system. Scaling AI accelerators is a long term capacity planning exercise, not an auto scaling group setting. 51:27 Speaker 1 The anti pattern 6 will have humans review the edge cases. 51:30 Speaker 2 Human review is not magic. It's a queue with finite throughput, high latency, high monetary cost, and sphere welfare consequences. If you invoke human review in a design, you must size it. What happens when traffic spikes 10X? 51:43 Speaker 1 If you look at that list, five of those six anti patterns are exactly the same error. 51:48 Speaker 2 They are. The error is naming A component instead of making a decision. Anyone can memorize the nouns Vector DB classifier fine tuning. A staff level engineer explains the why and the trade-offs behind the nouns. 52:02 Why the Core Triage Method Will Always Apply As we close this out, what is the real take away? Because honestly, the specific list of 10 questions we just went through, it's going to age. Question 10 about observability didn't even exist as a discipline 2 years ago. Next year, there will be something on this list that nobody is even thinking about today. 52:18 Speaker 2 The questions will change. The tech will change, but what will absolutely not age is the triage method. What is scarce? What does failure cost? How fresh must it be? If you sit down in that room and you run those three questions in the 1st 90 seconds, the architecture will literally start designing itself. 52:35 Speaker 1 And underlying all of that is the spine. Say it one last time. 52:38 Speaker 2 Name the scarce resource, then design around it. Think back to that candidate in the cold open. They knew every single fact required to pass that loop, but they failed because they didn't take 90 seconds to decide what the system was actually short of before they started drawing boxes. 52:55 Don't be that candidate. Say the scarce resource out loud. It cost you exactly 1 sentence and it changes the entire trajectory of the interview. 53:04 Speaker 1 All right, I promise some homework. Pick the three questions from this list that you at least want to be asked. It's probably question 7 on evils, question 9 on agents, and question 10 on observability. For each of those three, write out the triage answers. Three sentences each, Nine sentences total. 53:20 It will do more for your interview prep than grinding another 40 practice problems on the classic Canon. 53:25 Speaker 2 And for our next deep dive, we're taking one of those hard ones and going all the way down the evaluation harness. We are going to explore exactly how you test a system that gives a different answer every single time you ask it. We'll look at the mathematical flaws of LLM as a judge and why your aggregate quality metric is lying to you. 53:42 Speaker 1 Because at the end of the day, a system you can't measure isn't a system at all. It's just a hope. You aren't engineering, you're just praying the model does what you want. We'll see you next time.

Podcast Summary

Key Points:

  1. Classic web design fails AI system interviews because it wrongly assumes compute is elastic and cheap, while GPU capacity is constrained by memory bandwidth and long provisioning times.
  2. The core mental model is to name the scarce resource first and then design around it, since AI systems shift scarcity between GPU compute, token budgets, context windows, and human review capacity.
  3. A three-question triage process should be applied in the first 90 seconds
  4. The serving layer requires continuous batching, disaggregated prefill and decode, and two-phase quota commits because GPU utilization depends on dynamic scheduling rather than simple worker pools.
  5. The retrieval and memory layer must defend the finite context window through structure-aware chunking, hybrid retrieval, cross-encoder re-ranking, and categorized long-term memory.
  6. The trust layer needs layered evaluation using benchmark suites, LLM-as-judge with bias mitigations, online signals, and decoupled policy from model in safety pipelines.
  7. The control layer treats agents as durable workflows with checkpointing, idempotency keys, loop detection, context compaction, and full-chain observability.
  8. Six common anti-patterns reveal candidates who name components instead of making decisions, and the triage method will outlast any specific question.

Summary:

This deep dive examines ten AI system design problems asked in 2026 interview loops, framed through four architectural layers: serving, retrieval and memory, trust, and control. The central argument is that classic web architecture fails because it treats compute as elastic, while AI systems are constrained by scarce resources such as GPU memory bandwidth, token budgets, context windows, and human review capacity. The recommended approach is a 90-second triage process asking what is scarce, what failure costs, and what the freshness requirement is, followed by the core rule: name the scarce resource, then design around it.

In the serving layer, candidates must explain continuous batching, disaggregated prefill and decode, and two-phase commit quota systems. In retrieval, they must justify structure-aware chunking, hybrid search, cross-encoder re-ranking, and categorized memory with invalidation rules. The trust layer requires layered evaluation, LLM-as-judge mitigations, and decoupled policy from model.

The control layer demands durable agent workflows, idempotency, loop detection, and privacy-aware observability. The episode closes with six anti-patterns that signal shallow understanding, concluding that the triage method will outlast any specific technology or question.

FAQs

GPU capacity is not elastic on the timescales of web systems. Provisioning GPU nodes can take minutes to tens of minutes, and the capacity may not even physically exist in your cloud region at any price.

The KV cache, not the model weights. Model weights are static, but the KV cache footprint scales with batch size multiplied by sequence length, so long prompts batched together can exhaust VRAM before you hit any compute limit.

Embedding similarity can treat prompts with opposite meanings as near-duplicates. For example, 'patient has a history of high blood pressure' and 'patient has no history of high blood pressure' score around 0.98 cosine similarity but require completely different answers.

RRF combines ranked lists by position rather than raw scores, using the formula 1/(k + rank). Averaging is mathematically flawed because BM25 scores are unbounded while cosine similarity is bounded between 0 and 1, so BM25 would always dominate.

Run a semantic search on existing facts before storing a new one and prompt the extraction model to resolve conflicts, allowing recent extractions to overwrite older ones. Also build a user-editable memory UI so users can manually delete or edit facts.

Prefill processes the entire prompt in parallel and is compute-bound (FLOPS-heavy), while decode generates tokens one at a time and is memory-bandwidth-bound. Disaggregating them onto separate worker nodes prevents them from stalling each other.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.