330: Why Claude Gets Dumber—and How Lawyers Can Use AI Better
42m 16s
The podcast features Tommy Eberly, co-founder of Docket Drafter, discussing how AI can streamline legal document drafting, particularly for litigators. Tommy, a software developer, explains his journey from coding by hand to using AI tools like Claude Code, which he now uses for eight hours daily. His company, Docket Drafter, helps lawyers draft documents like complaints and motions inside cloud-based tools by using an MCP server that allows Claude to interact with a writing pad optimized for AI. This approach avoids the inefficiencies of Microsoft Word, which embeds formatting data that consumes up to ten times more tokens than plain text, slowing down AI and increasing costs. Tommy emphasizes that Word files are "the enemy" for AI agents, as they are complex and hard to edit without breaking formatting, whereas Docket Drafter uses an agent-native format that renders faithfully to Word. He also addresses the "context window" issue, explaining that AI models like Claude have a soft limit—around 50-60% of their token capacity—where intelligence drops sharply, causing hallucinations or forgetfulness. To combat this, he recommends starting new chats frequently and minimizing unnecessary context. Additionally, Tommy highlights that PDFs are similarly problematic for content editing, though Claude is good at page-level manipulation. Docket Drafter stores prior work for reuse, allowing lawyers to modify templates efficiently, and supports redlining for feedback. Overall, Tommy provides practical advice for lawyers to optimize AI use, reduce frustration, and improve accuracy by understanding token efficiency and context limits.
[MUSIC] All right, so welcome to the podcast and let's start with a cold open. Here's a sentence I did not expect to hear from a software developer. Quote, you should probably just run Microsoft Word's spell check. And the person who said that is my guest today, Tommy Eberly, who is the co-founder of a company called Docket Draftor, which let's litigators do their word document work from inside the cloud instead of fighting with word to try to make it talk to cloud in a different way. So he has been showing up in our AI sessions, which happened twice a week in intercircle. And he's been giving some really great advice to lawyers. The best really that the group has gotten. And so I wanted to invite him on the podcast to share some of what he's been sharing inside of our discussions. Tommy's not a lawyer. A year ago, he was writing computer code by hand. He spends most of his time, hours and hours a day inside of cloud code, not just for coding, but for marketing, customer research, and reading actual court filings of lawyers. And what makes it worth your time to listen to Tommy is he can explain what's happening with AI under the hood as best it can be explained. And once you understand what he is saying, a lot of your frustration around these tools will start to a make sense and hopefully diminish things like, well, why is Word? And eating usage limits much faster than if you had plain text. Why does cloud get noticeably dumber halfway through a long chat? And why is AI tool you're paying for give you six paragraphs when you ask for four sentences and why it's maybe not your proms fault? So we're going to get into all that and where he would push back on the legal tech industry, including his own corner of it. But first, Word from the sponsor or at least from me on behalf of the sponsor, Smith AI. So if you're thinking about hiring somebody to answer your phones, that might seem like the obvious next step, but then you start doing the math, payroll, training, HR paperwork, sick days, turnover, all that stuff. And that is why literally thousands of lawyers skip that route and use Smith AI instead. Their AI receptionist is they always on front desk answers every call around the clock, screen leads, blocks, bam calls, schedule appointments, and send call summaries straight to your phone, your email or your CRM. There's no onboarding. This is really fast. No shift changes, no stress. Smith AI has been at this for more than 10 years. They handled more than 20 million calls. They rated five stars on G2, trust pilot, and clutch. So if you're tired of hiring headaches, especially with receptionists, let Smith AI run your front desk. You can try it risk-free for 30 days. The link is in the show notes. All right, let's get back to our guest today, Tommy Eberley, who, as I said, is a software developer. I've explained his background, let's just dive right in, and welcome Tommy to the podcast. Welcome. Ernie, great to be here. Appreciate the invite. Yeah, no, I'm glad you agreed to do it because you've shared a lot of great information. With the lawyers in the inner circle, and I know we're going to share some here, so let's just dive right into how it is that you went from being software developer, coding by hand, then using AI, and then creating a product, which I know you will explain called Docket Raffer. So walk us through that preliminary journey. Yeah, so I started out at college as a software developer. I was actually working at a high frequency trading company in Chicago doing, they were doing options market-making. So these low-level trading systems, and after a few years, I quit that, and then I joined a Fintech AI startup. It's called Boosted AI. And I worked there for a few years. I actually moved to Austin, Texas during that time. And I don't know. Kind of always wanted to do my own thing, and decided after three years to quit, and I convinced my good friend from the University of Michigan, Grant Willard, to join me on this venture. We actually didn't have any idea what we were going to do when we got started. We were just doing general AI consulting. And after a little bit of just general consulting, we got connected with a guy who wanted to law firm, and we're doing some document automation stuff for him and started to get our hands a little dirty and legal. And then from there, we started going into the New York State court system, which we figured out was very open. You could read all the lawsuits. You could see what attorneys are filing and see their actual work products. And we found some guys that were doing hundreds of cases a month, and specifically guys that were representing defendants in consumer credit cases. So these guys are representing guys that are getting sued by JP Morgan, by velocity, investments, those kind of places. And we reached out to a few of them and convinced somebody to let us build a system that would generate their answers. So you upload a plane, and the answer gets spit out because they're often very similar. The affirmative defenses are the same. But from there, we built Docker Drafter and have been trying to expand. Now it's more focused on attorneys that are already working in Clawed. So we take here, this playbook we call it, which is your templates, your court formatting rules, your prior arguments. And we make that all available and really easy to use in Clawed. There was a Clawed code. We started with that. I think it came out in early 2025. And when it initially came out, it was really expensive because you had to use API credits. So I didn't actually use it. I was using cursor, which is kind of like the first edition of Clawed code, like the grandfather almost. And before that, I was just going into Chat GBT and copying and pasting code into my editor, which I think is something a lot of people probably still do. You paste into Chat GBT, ask it a question, then you copy it out into whatever tool you're using, whether it's Microsoft Word, PowerPoint, whatever. And then eventually they made Clawed code available on the subscription, which meant it was finally affordably. You could pay $100 a month and basically get unlimited usage. And from there, it went from a cool little experiment that Grant and I were doing until, okay, we're using this for every single thing. We're not even writing code by hand. And then from there, it became, well, we can do a lot of stuff with this that has nothing to do with software development. And my day right now is honestly like eight hours a day in these tools. And a lot of attorneys we talk to and honestly, just other people and other industries, they're getting into co-work, they're getting into these coding agents, which really are just general purpose knowledge worker agents. And once they get into them, they never want to go back. It's like this one way door. We haven't met some of those dove into Clawed code that has stopped. But we have not seen no person. Right. Well, and I was seeing that with the lawyers in my group, you know, we went from using Chat GBT and individual prompts and then got agentic stuff. And then Clawed code came out and Clawed code came out and it was a tsunami of people, you know, recommending and using it and so forth. And of course, inevitably the lawyers did too. And then some of them that were using co-work were then shifting it over to code. And you know, people find where they're most comfortable. But the idea that once you learn that you're doing a lot of work and maybe even most of your AI related work inside of this one area, you're like, well, why do I want to leave? You know, like I'm already here, I can query my email, I can, you know, have it do stuff for me. And so yeah, we have, as you know, because you've been meeting them and talking to them, some folks that are intensely into it. And so when, and of course I was watching them talk about, oh, there's a word connector for Clawed, but it really sucks. It's really, there's a lot of friction. It doesn't work great. And then when you talk about what you, what your product did, I said, ah, okay, well, this is a match made in heaven. So tell us a little bit about how you're, how a lawyer who's using Clawed code or co-work can use your service and then interact with word in a more efficient way. And why that can happen efficiently. Yeah. So when you're using DockerDraft, you're really just talking to Clawed and Clawed is going to interact with DockerDraft for on your behalf behind the scenes. So to get it set up, it's just an MCP server, which I think your audience will be familiar with at this point you log in. At that point, you just draft matters. And so that would be like complaints, answers, demand letters. We can even do forms like for, you know, if you're filling out court forms like in probate court. And you talk to Clawed, it talks to DockerDraft, and we are able to draft the documents in an agent friendly format, which makes it really easy for Clawed to do. It makes it very token efficient and fast. And then we had a way to take that agent native format and faithfully render it to a word doc. It works right. It's going to fill out the right fields in the form. And your agent doesn't have to fight with Microsoft Word, which as important as it is for lawyers, because it's really great at formatting and making sure like the caption has the exact right spacing, it's kind of one of the worst file formats ever to use for an agent. It's very, very difficult to work with if you're not literally using Microsoft Word. Mm-hmm. All right. So it's token efficient, which means the cost of using it is significantly less if you're doing it through DockerDraft. As far as the files that the lawyers would then be working with, when they use DockerDraft, where would those files be on their computer and your servers? Where would they be? Yeah. So the documents they're drafting are going to live on DockerDraft servers. And at any point, you can download a Word doc and see what it actually looks like. And you can go into the downloaded Word doc and actually redline it. And so, you know, this is something that we found people do. And this is kind of how it works in software development. You go, "Hey, quad. Modify the argument in.1 of this motion. I want to make it use this other case, for instance." And then you, it's work. And if you have any comments, you can just redline them.
the Word doc and upload it back into DocaTrapter and we'll be able to put that back into the system basically and incorporate the feedback. In terms of the input documents, so usually when somebody's writing a motion or an answer, they want to get caught contact. So what was the complaint about? If you interviewed the client and they told you information that's going to be used in the argument, that's relevant as well. That'll live in your computer or in Google Drive or wherever you prefer it. Clawed is going to go and talk and read those files. It's just going to use DocaTrapter as like it's writing pad basically. And so I would imagine that for a lot of these documents that are being drafted, the lawyers are going to use DocaTrapter during the drafting process and once it's been filed or dealt with, they're going to save it on their place, their server, and they're not going to necessarily need to keep it on your servers, is that how it works or what? Yeah. So that's totally correct. If you download it and if you have a place you like to put it like Google Drive or SharePoint or something, obviously you can do whatever you want with that. DocaTrapter by default, we store all of your prior work and we could obviously set it up if you don't want to store your prior work. But the benefit of doing that is if you work in a niche where the same arguments come over and over again, the same legal issues come up, it's often really useful to have your prior work available to Clawed because that way it's not regarding from scratch. So if you have a template or if you have a motion that you use and mainly just swap out the facts, you want that in the system because then it makes it really easy for Clawed to go in and go, "Oh, I'm just going to copy the old one, swap out a few things," which surprisingly we learned is a lot of attorneys that are doing volume in particular, they grab an old word doc of emotion and they just start from there and modify it. Right. And grabbing the word document is inefficient because as you were starting to say word is kind of the enemy in this process. So maybe let's get concrete and explain a little bit more about why it is that Clawed struggles so much with word files. Yeah. So Clawed and what I'm saying is Clawed specific, but this is going to apply to any large language model, any AI tool, so Chatubt, Gemini, all this is going to apply. Clawed operates on text and when you have a markdown file or you have even HTML or just plain text, that's the most efficient for Clawed. Because there's no or limited information about the formatting. You're not really specifying that much about the margins have to be precisely one inch away from the left or this heading is going to have slightly different before, after spacing. When you're in Microsoft Word, all of those settings you have in the styles pane and all of that formatting stuff that honestly is not always the easiest to do, even in Microsoft Word, that has to be embedded into the file somehow so that word can understand how to display it. Because that information is in there, it makes it really hard for Clawed to deal with because that's like 10 times more tokens than just the raw data. The information required to precisely specify the formatting is about 10 times more than the actual text itself. If you want to go and modify a word document, it's one thing for Clawed to change around a little bit of text, but if you have any kind of bigger edit, you're going to break the formatting and it doesn't really have a way to validate that it worked because it doesn't have the same vision that you or I do really. Right. Yeah. I mean, over the many, many years, it's been an existence. It's gone from a simple little one-celled organism to this very complex thing, which the user's not going to notice, I mean, well, it's complexity at the user interface level too, but the user's not going to see all that stuff, the hornets nest of cruffed underneath. They just know, okay, if I ask Clawed to do something, why is it taken so long? And that same problem comes up with PDFs, right? Because PDFs are also, even though seemingly they're very simple, they just display the document the way it would be if you printed it out. There's a lot of stuff happening behind the scenes that Clawed or whichever one has to wade through to do anything useful with. So you're burning through tokens and that's inefficient. If you want to be, you say you're in high frequency trading earlier and then you went into work with these lawyers, they're like high efficiency litigators, right? And you can't get bogged down with cruffed if that's what you want to operate. Yeah. And you mentioned a comment I had said earlier about when you get to halfway into the chat, half of the context window, the intelligence starts to drop. A lot of people now, they're worried about token costs. It's going to be so expensive if I have to use all these tokens. It's totally true, but more importantly and probably more of a pain is the more tokens use the dumber the models are. If you can get by with fewer tokens, the intelligence is literally going to be higher than if you are 80% of your context window full. And that's totally right. As for the PDFs, they seem really simple because hey, it looks the same on my phone as it does on Ernie's laptop, as it does on my laptop. The reason they're that simple is because under the hood, it's kind of a disaster file format to be honest. They're very hard to edit. But one thing I will say about PDFs is, and what Claude is very, very good is, if you need to splice and dice PDFs, so removing pages or smashing things together, it's excellent at that. And I would totally recommend it. But if you need it to modify the content of a PDF or fill out a form, it can kind of do it. But depending on the complexity, it's probably going to be a little bit painful. All right. So I want to mention that one of the things I'm going to give a link to is your YouTube videos because they're excellent. You do a really great job of explaining things very clearly, very simply and not just theoretical things, but things that have a practical utility if you understand them. And your video about why it is that Claude gets dumber, which you shared inside the inner circle, which is great. I wonder if you would mind going through kind of the overview of that because that's a really important thing. I've talked about that in the inner circle a lot because it's a really important thing to understand why it gets dumber and also to know you don't want to drive it to the end of the road and then get out swap cars. You want to like, you know, get out a little sooner and you know what I'm talking about, but you explain all that. Yeah. So I'll start with context, which I think your audience is probably familiar with, but just to give an overview. When we're talking about the context window, if you're thinking about a Claude or a chat you beat each app, every piece of text that you put in, every file you upload, every file it reads, and every piece of text or image that the AI generates goes into the context window and operates space in AI. It's not words, it's tokens and a token is not really a word, but basically one word of English. It's about, I think it's about three words of English is about four tokens, give or take. And so basically the more text, the more stuff you have in there, the more tokens you're actually using with every new token that gets generated. Each model is going to have a token limit. Usually it's around 256,000, maybe a million depending on the model. The numbers aren't that important, but if you go over that limit, that's what I call the hard limit. It will literally crash. It will not be able to physically output another piece of text. And so you can never go over the million or the 256K depending on the model. However, there's a more real limit, which is far earlier than the 256K or the million, that's the soft limit or like the intelligence drop off. And it depends on the model, but it's usually around like half or 60% full or even less. Once you approach that point, excuse me, as you approach that point, the intelligence is going to slowly decline. And you're going to notice that it starts to maybe hallucinate a little more, forget things you got reminded of stuff. Once you cross that 50 or 60% thresholds, that would be 128K tokens or 500K depending on the model, you're going to see a giant drop all at once in intelligence. And that's what you really want to avoid because it's not linear. It's like very sudden at that point. And you'll probably realize when you're there, when you start to get really frustrated with it and it starts to not understand what you're saying or contradicts prior instructions. But that's kind of like the high level overview. And like the simplest way to get around it is just start new chats a lot more than you like to. And people find it annoying and I agree, but that's how you're going to get the best accuracy is starting your chats and being very thoughtful about the context that you're providing basically. So that being true and then thinking about how nowadays magically you can give these models some amazing work to do. Like I gave, I've been giving mine these projects from like, yeah, here's a whole bunch of stuff like sift through it and just process it. Well, it's acting on its own. So it's running at night, let's say. So if it's running at night, you know, there's a point where one, it'll hit the, it'll hit the, you know, the hour usage or whatever and then it'll start drawing down from my, you know, reserve tokens. But at some point, I would guess it's going to hit that drop off threshold without me knowing it in the middle of the night. How does one deal with that when you're like, I mean, autonomously depends a lot on the task. And I will say the models, like they're getting better at handling the longer context. So if you're having it do something like, you know, go through the internet and provide you some links of a topic you're interested in, then maybe it's fine because if there's some false positives, like if it gives you a link and you go, Oh, I don't really care about that one actually, not really that big of a deal. Like it's okay. If it's 80 or 90% correct and it messes up a little bit, if the stakes are higher though and especially viewing something like a legal document, that's where you might not be able to tolerate that level of error. And so your options would probably be to scope it down as much as you can and try to prevent it from needing that many tokens in the first place, which I know is kind of abstract and really depends on the example.
But generally I would avoid overnight like long running agents that you're not gonna be babysitting for stuff That's missing for it or like you have a very low margin for error If it's more exploratory or something where you can tolerate some nonsense, then it's totally fine It's more just to get a feel for where it's good and where it's not basically Yeah, I mean, I'll say explicitly what one of the projects was since it relates to this podcast And I was kind of shocked at how easy and fast this was but it fast, you know, relatively speaking was I said, okay I need to or I want to go get all the recordings of these you know 325 episodes prior and Transcribe them and then sift through the transcripts and pull out the highlights, you know just Interpret them because the raw transcripts generated from the file and They were a little messy right so I found and this is this is almost vibe coding. I just went to it said, yeah, that's what I want to do What's the most efficient way to do it? It's a well. There's just tool called whisper We can run it in the terminal we get downloaded and run in and it calculated Based on the power of the computer I was using how long it would take and you know, yes, let's break it down into segments and All that kind of thing so it seemed like that task Was very pedestrian for it like you know, it wasn't a lot of hallucination wasn't gonna be a problem But what came out when it started interpreting I was shocked it pulled out all these insights because it's looking at You know 90 hours. I don't know a lot of hours 100 hours of prior episodes and I was shocked at how much useful information was to me and then I asked to analyze going forward podcasts and tell me what topics to suggest So it's you know when it when you let it look at a lot of data and it can do it effectively It's kind of shocking what comes out of that process Yeah, and I will say you said it's kind of vibe coding that is vibe coding Ernie It doesn't have to be holding it sass app, you know, it's writing many programs A custom Ernie program to pull out the transcript the other commonal make is it's not just about the length of time the agent is running Because what really matters is how much text is it emitting the entire process and if you think about it like if there was a website Where you had to dump a MP4 file and it would give you a transcript that might take a few minutes to run It's the same kind of thing when that's happening on your laptop and while the program is running it is not using any tokens It's just waiting for that to finish and so even if that might have taken hours to pull it out of the transcript That probably didn't use very many tokens because it just had to run that program You know a thousand times and that's where a lot of the time was and that's the other distinction Yeah, that's true It didn't use a lot of tokens and it was just a question of the computer and was was funny is It was going to take it like you know seven nights to do it all and at some point I thought you know Why don't I just use my Mac mini instead of my laptop for this and the Mac mini's got an M4 chip You know like it's twice as powerful as the laptop and It's designed to be on all night because that was another thing that the vibe coding Turned up was it said oh, you know your computer keep shutting down We need you need this caffeinate thing whatever. I don't know like caffeinate when I find bringing home Install it and it installed it and that kept the computer on Which it was plugged in but it would go to sleep because it was a laptop Whereas the mini all of a sudden we're going at 4x the speed and processing and that was faster and more efficient But I just I was struck by the fact I was working in the terminal because that's where it could install things and work more efficiently And I'm not used to do that. I mean I could do it But you know it looks like a lot of goblies are good to me and I had to install things and as you said It you know it was Five coatings it was telling me yeah, you need to install this just run this pseudo command or whatever in unix And I was doing all these things and it was warning me when you know I shouldn't do you know Should be careful and pay attention, but it did everything Simply by me copying and pasting the instructions. I would like run it over there and then say okay This is what it said. It's oh no, we messed up that. Let's fix this. So that's where You're right vibe coatings happening when it's creating a little mini programs But even like when you're telling it to do stuff and it's you know putting out code Which I don't know what the hell that shit means But you know I can paste it in and it will interpret it and tell me what why it worked or you know where I need something else So I use this analogy earlier, but it really does feel like slipping on a banana peel Except instead of hurting yourself you come out with something useful and It's crazy people you heard today somebody describing how they Just created this whole thing that they didn't even intend to start creating Because that's what you can do now Yeah, and I'll say just for the audience. I don't know how familiar there are with when you're talking about the terminal If you've ever watched a movie where there's like a hacker on the computer and they're on this like dark screen and a bunch of text is flying by and they're just hammering in Commands that's the terminal and you can actually use a computer without meeting Anything other than a keyboard and a command line. So you're literally typing in text specifying exactly what to do You're telling it copy this file from this directory in and this other directory Go download this website run this program and yeah, it's gonna look like go to most people that are you know experience programmers But the thing with the terminal is you can do literally anything you want with it Which is really scary if you don't know what you're doing But if you have clot or another AI agent that's really powerful because it's very expressive and that's where if you get all a freedom anytime and quad Co-worker quad code does anything for you. It's actually running a lot of terminal commands under the hood. That's how it does anything All right, so we start talking about using word and of course, I'm realizing that if we don't talk about track changes There might be a lynch mob, you know, assemble because lawyers who use word want to do track changes is that something That your your product can help with or clog can help with yeah So in docket draft your word drafting in this agent native token efficient format and we track version history for every single thing That's happening and we have the agent actually explain why it made every change so you can go back and I'm you can try out a new argument and throw it away The agent will remember and it'll have notes about why it did what it did and anytime you want to compare two versions We can generate a red line so if you're like hey, I don't know Maybe your associate was working on docket draft here and made some edits and now you have the latest version You can say give me the red line version and it'll actually pop it out with everything highlighted and explain to you All right, well, that's good and very comforting um The other thing that you helped us understand Was something that comes up I think I think you can do this in co-work now, but in code for sure In the middle of a chat you can switch models so you can start off saying okay, this is Low level I'm going to use on it. Oh wait now. This is thinking I need opus um what same concept in chat GPT and so understanding When would you use the lower level one or the middle level one like you know high-coup is like the most basic Then sonnet then opus then fable if you have access to fables. So Give us a sense of when you would use those different models Yeah, and I'll start by explaining that the two axes that we have here one is the model So that's high-coup Sonnet opus fable and then the other is the thinking level and so that's going to be low medium high extra high ultra high Whatever they're naming it those are two distinct things and also point out that we're talking about the anthropic family of models Open AI has their own family and so does Gemini and and they have these different axes Hi-Q is the smallest brain like literally it is a smaller model. It's not as big as sonnet Sonnet is smaller than opus opus is smaller than fable and when I say smaller You could kind of think of these as like a brain with neurons and fable has the most neurons So it has the most potential for intelligence The as you go more neurons you're gonna generally get better intelligence They're gonna work better at longer contacts and they're gonna hallucinate less When you're using a particular model the thinking level is also gonna affect the perceived intelligence But the trade-off there is the higher the thinking level the slower it's gonna take in the more tokens It's going to eat when a model thinking It's actually emitting tokens, but it's kind of like if you were to just talk to yourself in your room and maybe you're talking about an argument like Oh, well, I just added this one case citation and it seems really good But let me think about it again. Let me reread this section and cross check it and it kind of sounds like nonsense But that's literally what the models are doing. They're just emitting more text. You can't really see it the model providers Don't really let you see the full thinking traces Generally if I'm using cloud I tend to stick with opus mainly because fables expensive and it also a little bit slow Sana is okay and for some lower level tasks you can get away with it And if you're concerned about usage or tokens, I would play around with it Especially on the lower thinking levels. I tend to prefer Lower thinking because I'm a little impatient and I kind of like to be more the babysitter and hands on the keyboard So I don't want that they're in weight I would generally avoid hiku. It's not very good and other than maybe using it for Like converting PDFs to text or some other very limited use cases Sana is just way better and your usage that fast. So I would just generally avoid hiku Fable is great. Honestly, the big thing is just it's slow and you might eat your limits a lot faster than you'd like Right Yeah
I don't use Hiku. I have noticed in one of the tools I use, which is slashy, this email service. It defaults to Hiku, and so I didn't notice that it was defaulting, and then I realized that it does a fine job on the basic emails that it cranks out. There's not a lot of thinking involved. It's more about speed and efficiency. And so it seems to work there, but I don't choose it when I'm using cloud. I do notice, and I try to pay attention to the expressions of thinking. And we were talking about this today earlier in our meeting that, like when I asked for a recommendation for binoculars, and I was like, all right, I gotta buy some binoculars because after replacing ones I gave away, I mostly use those at Saints Games, which we still have Saints tickets will drive to the games from Panama City Beach and go see the games, but I got used to like being able to look at what I wanted to look at, so I need binoculars for that. I also have a lot of birds around here when I look at the birds, and then maybe I wanna do stargazing. So those are three different uses. And so I figured, well, I'll just ask Opus because I wanted it to think this through what binoculars should I get? You know, give me three recommendations. And when I was watching it go through its analysis, it, at some point, said, oh, superdome, let's check the bag policy and whether you can bring in a binocular case. And I had never been bringing in a binocular case. I didn't know you were prohibitive from doing so, but it makes sense. So it did a whole analysis of that. And I went back and I looked at what the analysis was 'cause you can do that. And I saw like, wow, I'd really chewed through that. And I assumed it if I had chosen Haiku, that would not have happened, right? - Yeah, exactly. And it's hard to say for a specific task, like it might be the case that Haiku found some of the same stuff as Opus or Fable, but generally speaking, the higher tier the model, the more analysis it's gonna do, the more of the edge cases like the stadium bag policy, it's gonna think of a head of time. And the higher the thinking, that's also gonna have an effect on its thoroughness. And you know, if you look at something like just buying binoculars, you know, is it the worst thing in the world if it maybe misses one of the criteria? Like maybe, maybe not. It kind of depends on what you're using him for. - Right. - Because you look legal research and you're saying, hey, I wanna site check this one opinion in this brief and you're actually trust the model's output. That's where maybe it makes more sense to use a higher tier model and put it on higher thinking because you want it to get it right and you're okay with it and paying for that because the stakes are higher. - Right. - All right. I think there was another thing that you, I'm trying to remember with that question I had, do you create a video for about the-- - Was it the job compassion, I believe. - Yes, yes, yes. Well, I'll just do that. - Yeah, so I can pass you. - Talking about the context window and again, to remind everybody, there's a hard cutoff. It's either, you know, it's 256,000 tokens. Sometimes it's a million, it depends on the model. If you get over that limit, everything breaks. Like they literally cannot output any more text and if you're using quad or chat, you be T and you reach this point where it just gives up and completely crashes, you could imagine that would maybe annoy you and that wouldn't be a great experience. And so what, andthropic and OpenAI do now is they detect your context of usage. And once you reach a certain point, usually it's around 80, 85, maybe 90%. What they're gonna detect is, okay, well, we're really close to that limit. Like we're almost out of gas here. What we're gonna do is take the entire chat, have a different quad, summarize the chat and turn it into a set of notes. So the whole chat's gonna have a giant conversation. Like, hey, quad, read through this word doc and do this legal research. The quad's gonna say, hey, here's what I found about this case. Really, really verbose. When you are after compaction, that chat literally gets replaced with a set of notes. So it'll say, like, user asks me to do legal research in the state of New York. User and I drafted this word document. And so it still knows generally what the session was about. And it's still able to pick up kind of remembering everything, but because it's a summary and because it's just the notes, it's not necessarily gonna get every single thing there. Like if you took notes on a meeting, that's not the same as going to the meeting and listening to every word. You might get the general idea, but you might miss some of the really important details. And so I think we'll find after compaction, they'll get what seemed to be hallucinations, 'cause it's like, hey, Claude, we already talked about that. It was not really a hallucination. It's that Claude generated notes about your previous conversation. And the info that you're asking about is literally not in the notes. So it has no way of knowing that. - Hmm. All right, let's talk about files and folders. 'Cause you have said this a couple of different times. And of course, I bring it up as well. 'Cause I strongly believe that one of the, like simple, powerful things that not enough people are doing is letting Claude manage files and folders. For one thing, you have files and folders, you have info in there. You know, lawyers learn that they can point it at those files and folders and that's fine. But you can also tell it to like evaluate the naming convention. You should things be moved around. And that's what I did. I had Claude go through both my business and my personal drop-ox and Google Drive and give me recommendations and then just say, okay, fine, those are good recommendations. Just execute them. So files and folders, pretty powerful, right? - It's shocking. It's such a simple concept. Like, oh, of course, you have files and folders. Like people have had that on computers probably since they came out. It's very powerful to understand that concept. And I would encourage anybody, like if you're anything like me, your downloads folder is a complete mess. Like you have, you know, it's years old. You've got all this crap in there. A lot of stuff that you probably should delete that's just eating up space. Have Claude just start a chat with, excuse me, coworker Claude, tell it to go through your downloads folder and just evaluate everything. And all you have to tell it is, hey, my downloads folder is a mess. How can we clean this? And even that little information, it's able to see every file in the folder. It can read the file names of everything. It can poke and prod at the different files. So if it sees a file that's like some PDF that has a completely garbled name, it can read the PDF and go, oh, this is actually that complaint you downloaded from the New York State Core website. The name just doesn't really identify it. And once it does that, it'll say, well, if I can read the folders and files, I can also write to them. I can modify their structure. I can create new folders. I can move things around. And so if you ever wanted to organize like a desktop or something, that's a mess. Claude can do that in a breeze. And it's really, really useful. - Yeah, and it'll find both files and folders that contain, well, the folders contain nothing that you accidentally created or something. Or files that are named untitled because there's nothing in them. And for you to go hoover around and find that stuff, it'll say, I found these. What do you want me to do with them? And the answer is delete them. And then you approve it. And attorney I was working with wanted to start doing this and he had, he was looking at, he has just a whole slew of matters. And he's like, well, this is great. I can kind of organize these things. And I do one kind of work and my partner does another kind of work and we could clean it up to where it's more designated what the files, where they belong. But we could also say, hey, some of these files are closed. They're not active. Let's move them aside. And we talked about this and I assume that this is valid. I said, well, I'm sure it can look at a file folder and say, nothing new is gone in there in like three years. That's probably a closed matter. And of course, you approve it and you go, it's fine, now move it to archive. So it can do that, right? It can move through those things pretty quickly. Yeah, so if you open a file on Mac OS or Windows and you right click it, there's metadata on the file. So data about the file and you'll often find stuff like created at last modified time, who created it. And because you know, you or I can see that, well, Claud can also get all the metadata. And so it can easily figure out like, okay, well, this file you downloaded yesterday, this folder you haven't touched in four years, at least according to the metadata. And it already points out, as long as you're not too brazen in your prompting, as long as you don't tell you to go completely nuts, if it's unsure about something, it's going to ask you and if it is going to do a destructive operation like delete or a move, it's probably going to ask you. And if you really want to be safe, just be explicit in your prompt and say, don't move, delete or rename anything before running it by me. And that way it'll be able to suggest without completely blowing stuff up and losing any data. Okay, and then one other case example that we discussed with one of the attorneys was he was having trouble. Oh, this is a recent one you'll remember. Getting Claud to not do something. He tell it, don't do that. Like, and you said it's the pink elephant problem. Could you describe what's going on there? - Yeah, so there's a term, I think it's from some psychologists, but the expression is don't think of a pink elephant. And when you say that out loud, obviously whoever's listening to you and their brand immediately thinks of a pink elephant, even though you said don't do it. And it's kind of a useful analogy with Claud because anything that's in the chat will be used against you. Anything you say can and will be used against you. And where this gets kind of problematic for people is when it goes onto like a tangent, or when you're working through an argument or having to do something and it starts behaving really poorly, it's not able to get that out of its brain. And if you tell it, oh, well, don't actually do that research that we did earlier because you did a really bad job, you're better off starting a new chat because it's still thinking about that. It has no way to not think about that kind of thing. - So yeah, let's have you close out by talking about a docket drafter. What would you like to see happen if Flora is interested in exploring this? Like what can they do next to work with you or find out more about it? - Yeah, so they could definitely go to our YouTube channel. We've got a lot of practical Claud tips. We've also recently released a bunch of free legal research plugins for Claud. And I have a video on our YouTube [BLANK_AUDIO]
explains how to install them, but we've got a court listener integration. We've also downloaded various jurisdiction statutes. So we got the entire US code available for download. It's index, very usable and cloud co-work. We've got the entire New York statute list. We've also got Florida. We've got federal court rules, and we're kind of adding to this. So I would say definitely check that out if you're getting into cloud. If you want to work with us, you can reach out directly to me. We have a link on our website. I've got my email all over the internet and on our website as well. You can reach out directly. And how it works with us is we basically figure out what jurisdiction are you in, what types of documents are going to be preparing. We get your workspace set up. And then when you log into DocketRactor, you can just go right in and say, hey, draft me a complaint for the Southern District of New York. And it's going to look right. And all the while you're using it, all those matters are saved. That becomes your firm memory. And it becomes something that compounds over time. And if you've got templates that you've liked already, you can send us those as a word doc. And we'll build it in so that the system already knows about your go-to answer or motion template, basically. That's great. And of course, obviously you'll help people understand how to use it and coach them. Yes. And that comes with one-on-one training with myself and Grant. And so we'll sit down. There's no stupid questions. We'll go over everything in AI. We can talk about the basics of cloud. It's going to work better if you're competent in cloud and understand some of these concepts. So we're very happy to make sure you're comfortable with prompting. And understand how to good AI intuition, basically. Right. Well, great. Well, that's some good info. I will put links to your email address and so forth. And everything else so people can see that. And your YouTube videos, which are exceptionally good. But thank you so much for appearing. And thank you for helping the lawyers out. You've been a wonderful resource for them. More of them show up more often and more eagerly because they hope you'll attend. And Nick, I know you met Nick is the same thing. Nick's been on the podcast. You guys have just been really wonderful. So I really appreciate it. Well, thanks, Arnie. And yeah, it's really interesting to see what lawyers say about these AI tools, where they get stuck, and kind of their understanding of it. So it's really awesome for us, too. So I appreciate it. Yeah, you're welcome.
Podcast Summary
Key Points:
Tommy Eberly, co-founder of Docket Drafter, transitioned from software development to legal tech, focusing on AI-driven document drafting for lawyers.
Docket Drafter integrates with Claude via an MCP server, allowing AI to draft legal documents in an agent-friendly format and render them as Word docs efficiently.
Microsoft Word files are token-inefficient for AI models due to embedded formatting, making them slower and costlier to process; Docket Drafter bypasses this issue.
AI models like Claude have a "soft limit" in their context window—around 50-60% full—where intelligence drops sharply, leading to hallucinations or forgotten instructions.
Starting new chats frequently and minimizing token usage improves AI accuracy and performance, rather than pushing to the context limit.
Docket Drafter stores prior work on its servers, enabling reuse of templates and arguments, but users can download Word docs for redlining and store files externally.
PDFs are also problematic for AI editing due to complex underlying structures, though Claude excels at splicing or combining PDF pages.
Summary:
The podcast features Tommy Eberly, co-founder of Docket Drafter, discussing how AI can streamline legal document drafting, particularly for litigators. Tommy, a software developer, explains his journey from coding by hand to using AI tools like Claude Code, which he now uses for eight hours daily. His company, Docket Drafter, helps lawyers draft documents like complaints and motions inside cloud-based tools by using an MCP server that allows Claude to interact with a writing pad optimized for AI.
This approach avoids the inefficiencies of Microsoft Word, which embeds formatting data that consumes up to ten times more tokens than plain text, slowing down AI and increasing costs. Tommy emphasizes that Word files are "the enemy" for AI agents, as they are complex and hard to edit without breaking formatting, whereas Docket Drafter uses an agent-native format that renders faithfully to Word. He also addresses the "context window" issue, explaining that AI models like Claude have a soft limit—around 50-60% of their token capacity—where intelligence drops sharply, causing hallucinations or forgetfulness.
To combat this, he recommends starting new chats frequently and minimizing unnecessary context. Additionally, Tommy highlights that PDFs are similarly problematic for content editing, though Claude is good at page-level manipulation. Docket Drafter stores prior work for reuse, allowing lawyers to modify templates efficiently, and supports redlining for feedback.
Overall, Tommy provides practical advice for lawyers to optimize AI use, reduce frustration, and improve accuracy by understanding token efficiency and context limits.
FAQs
Tommy Eberly is a software developer and co-founder of Docket Drafter. He started in high-frequency trading, worked at a fintech AI startup, and now helps litigators work with documents in the cloud.
Docket Drafter is a tool that lets lawyers draft legal documents like complaints and motions through Claude, using an MCP server. It formats documents in an agent-friendly way, making it token-efficient and fast, and can render them into Word docs.
AI models operate on text, and Word files contain extensive formatting data that increases token usage by about 10 times. This makes it inefficient and harder for AI to edit without breaking formatting, unlike plain text or markdown.
Lawyers can upload input documents like complaints or client info, and Claude interacts with Docket Drafter to draft documents. They can download Word docs for redlining, upload feedback, and store prior work for reuse in similar cases.
The context window is the space for all text and files the AI processes, measured in tokens. As it fills up, AI intelligence declines, with a significant drop at around 50-60% capacity, leading to more hallucinations and errors.
Start new chats frequently instead of continuing long ones. This keeps the context window less full, maintaining higher intelligence and accuracy, as the model performs better with fewer tokens.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.