The Case for Deterministic AI with Logical Intelligence CSO Patrick Hillmann
38m 46s
Patrick Killman, Chief Strategy Officer and COO at Logical Intelligence, shares his unconventional journey from crisis management to AI, emphasizing the need for deterministic systems in critical infrastructure. He explains that LLMs, while excellent for natural interaction, are probabilistic and inherently unreliable for high-stakes domains like energy grids, healthcare, or manufacturing, where even a 0.1% error rate is unacceptable. Logical Intelligence addresses this by building energy-based reasoning models (EBMs), which use physics-inspired energy landscapes to find optimal, low-energy solutions based on constraints, rather than navigating token-by-token like LLMs. This approach is more efficient and verifiable, as demonstrated by their strong performance on formal benchmarks like Putnam (98%), Verisoft (98%), and Verina (100%). Killman envisions a future AI "sandwich" where LLMs serve as user interfaces, deterministic reasoning layers handle core logic, and world models provide contextual data, enabling AI to operate safely in critical systems. He stresses that formal verification is essential for making failures visible and ensuring trust, especially as AI expands into areas where lives depend on accuracy. The company's team, including a Fields Medalist and veteran engineers, combines theoretical rigor with practical software development to push this vision forward, positioning Logical Intelligence as a key player in the next phase of AI evolution beyond probabilistic models.
Today on Shift AI, we're talking to Patrick Killman, Chief Strategy Officer and Chief Operating Officer in Logical Intelligence, a company building deterministic energy-based reasoning models for the critical systems that probabilistic LLMs were never built around, things like power grids, healthcare, and advanced manufacturing. Patrick sits at the rare intersection of crisis management in Frontier AI Strategy, having spent his career inside some of the biggest crises in modern business. From running crisis communications and cybersecurity in Edelman, to steering Binance through its most turbulent years, and a major DOJ settlement. We talked to him about why critical infrastructure can't be built on probabilistic AI, what an energy-based reasoning model actually is, and where formal verification is headed next. Shift AI is a weekly podcast and video series about the shift happening inside enterprises and how they're approaching AI and the Frontier of DeepTek. We cover the strategy, infrastructure, data, and security necessary to move quickly at scale. Every week we sit down with the leaders navigating this transformation in real time. If you want to learn more about sponsoring the show or about the Shift AI ecosystem, send us a note to sponsors at shiftdi.fm, or check us out at shiftdi.fm, where you can read about the episodes and all the other things we have growing up. With that, let's dive into the episode with Patrick Kilman from logical intelligence. All right, Patrick, thank you so much for being on the Shift AI podcast. I'm really excited to talk to you today. Thanks for having me on. So Patrick, I want you to describe your journey from where you started to where you are now at logical intelligence, just to give people an idea about what you've kind of been through from a career standpoint, and then kind of your position now in the company. Yeah, I'm going to guess I have a pretty random one compared to most of your guests. So most people probably know me from the crisis management space. So I came out of grad school, got a master's, economic policy, ended up falling into a very weird space. I was working at the European Commission. Russian Vaged Georgia happened in August of '08, and everyone in Europe basically goes on vacation to their shadow in the South of France, and I unfortunately did not have one being a grad student. So I got shipped off to the front to kind of deal with the policy and communications work like through the conflict. And so I fell into this weird space, working with large governments and corporations across the globe, and they had like really big overarching issues there trying to advocate. Whether that was GDPR, and how do you establish good policy around security and privacy to be a company, and how you navigate that to global scale, to more traditional conflicts. But then also cybersecurity. I worked at a firm called Edelman, which helped take multiple companies through the really bad years of ransomware to some of the largest ones on record, ended up working in-house at General Electric, trying to help them sort of navigate a series of issues globally, particularly helping people understand like the complex issues around advanced manufacturing and offshoring, and what that means for the future of work and for the future of the American workspace. Then went to a firm called Edelman After That. At Edelman, I ran their global crisis communications and their cyber security practice. I built the largest disinformation practice, while I was there with Cambridge University, with Dr. Sandra Vanderlinden, which was the largest at the time. And then I got a phone call from a company called Binance, which at the time was the-- and still is the largest crypto exchange in the world. They had found out that they had this massive DOJ investigation. They were struggling to help people understand like how their tech works. What does it mean to be a promotionalist network? And also then how to just pivot as a business and grow. And so I took Binance through the three most turbulent years in the history of the crypto space, helping them get settled, get through their cell with the DOJ, and help people understand like what is crypto? I left crypto, and I fell into AI, where I was introduced to a founder named Eve Badea. And after she told me what she had built, I literally called her the next day, and begged her to let me come and work for her. And I've been a firm called Logical Intelligence ever since, where I now am their chief strategy officer and chief operating officer. I want to get really into logical intelligence and talk a little bit about what you guys are doing. But before that, I always like to ask a question that grounds the audience and the mid-guest. And that is, what was your first job where you actually got paid? Someone gave me money for doing the work. Actually I paid, yeah. So my first job was a summer job in high school. I worked for UPS. I worked in their warehouse. There's basically two different tracks you can take when you're a high school working at UPS. You can either load trucks, where you load about one truck a day, or you can unload trucks, where you unload like 15 trucks a day. And I was fortunate enough to be able to go and unload 15 UPS trucks a day in 100 degrees summers. It's great. But anything do you learn from that that you take away, that you still kind of think about now? Honestly, I did. One of my biggest key learnings from UPS, which I've carried with me through life, and honestly really kind of helped me establish myself in the business world and helping companies navigate like really hard issues, is that you can have really complex manufacturing processes and systems and machinery. But in the end, there's almost always just somebody sitting their weapon boxes off the back of a truck. And so how we evolve from those systems, while keeping the average person in a role where they can continue to work and provide value in this increasingly complex industrial society, I think it's just really important and interesting. Yeah, absolutely. Well, let's dive in and talk about logical intelligences. There's so many interesting things about what you guys are doing. Talk about the players, because these are big time folks that have joined this company. And then also, what problem are you trying to attack right now? Yeah, so let me actually start with the problem, if you don't mind. Yeah, please. I wouldn't even call it a problem. So the way we see it is that AI is really in its first phase. And in this phase, we have LLMs, which are this incredible interface for the average person to be able to go and engage with an artificial intelligence. It's really good at talking and helping people to understand or simplify kind of complex theories. You can talk with that and it sounds really smart and it's very confident and it allows you to learn things really quickly that would have taken years sometimes in the past. Yeah. The problem is those as you know, in most of your viewers, know is that these are probabilistic systems, right? Which means that they are essentially just really good guessing machines. And that's not a knock. That's a feature, not a bug, right? It allows for a very simple and very natural sort of flow of communication between humans and machines. But as we've seen, having these problems in systems in place also creates an environment where AI is fundamentally sort of prevented from being an operate in some of the most critical systems we have on this on the globe. And that ranges from energy to transportation of defense, the semi-cultured space, healthcare. You can't have systems that are probably right because they sound good operating on you in a surgical field. You can't have probabilistic systems driving your car. Even if like you had an LLM that was 99.9% right, you wouldn't put your kid into a car that one in a thousand times is going to drive to the wrong location. And so for us, we saw that gap in the market and realized that in order for AI to really be able to permeate throughout society, we have to have a much more robust AI ecosystem. That's not to say that LLMs are going to disappear. They're not going to disappear. In fact, LLMs are really important. And as we start to have new architectures come to the marketplace, they're actually going to increase in their importance because they'll be able to have a role in the other areas that you can't really where they can't right now. But we need a new layer that is more deterministic that is going to provide a layer of certainty around AI so that we can actually start to optimize some of our more advanced manufacturing systems. And so for us, that's really what we wanted to build. To do that, you need a couple of ingredients. You know, building new architectures is like really challenging. So at our company, we have everyone from a fields metalist and Michael Friedman, which is for those who don't know, they give fields metal down like every decade. I mean, you'd say it's basically the Nobel Prize from mathematics but they give out a Nobel Prize every year, like, whatever. I can't walk down the street, I'm bumping into Nobel LLLR at these days but the fields metalist, not that often. I'm kidding, obviously. And you also have Jan the Koon, who is helping to direct our research and technology. You know, we meet with Jan multiple times a month and Jan helps us take a look at how we're scaling and how we're evolving our internal AI intelligence systems. And making sure that we're growing in a way that's going to help us break into some of these industries and overcome obstacles as we arrive to them. And underneath that, you have a team of engineers that range from 10, 20 years of experience, shipping systems, that met Google, Cruz, etc. To also international, a math Olympian champions. A couple of them actually have, I think, two championships globally each. And so you have this really unique genetic makeup, a really experienced software engineers who know how to build and ship products. And then you have the theoretical side that's able to help us, you know, take the most foundational and challenging elements of mathematics and apply that to software engineering. Yeah, absolutely. It's such an interesting challenge. And I know that the Jan has been talking about this a long time. And when I read some of the things he spoke about over the years, it's not the first time he's spoken about LLLNs. And it's still,
It seems like at least in this category of deterministic and critical infrastructure, companies that really matters, you're kind of making a bet against LLM's in that category. Would you agree with that? - Look, Jan gets up in the morning and forgets more about artificial intelligence architecture on his way to the bathroom than I will amass in my lifetime. - I think that being said, I actually don't think that we're a hedge against LLM's. I think we're a hedge against the kind of narrow perspective today that LLM's alone will be a unitary architecture. - Yes. - For a while we had this grand vision where all the major AI labs were kind of race and to create Jarvis or hell from 2000 on Space Odyssey. And I think that has disappeared. We're not talking about AGI like we were a couple years ago. And that's largely because this uncanny valley has sort of formed in the minds of most power users, but even like everyday users and engaging with LLMs. In that you start to view, I look at LLMs like a really confident intern. They're very quick, they sound really smart. And when they deliver you an answer, they deliver it with conviction and confidence. The problem is if you know the content really well, you understand that it's wrong in subtle ways. But because they're public systems, the more they start to have cracks in the foundations, as it starts to go into deeper and deeper levels of theory and thinking, it starts to get more and more and more inaccurate. And that's obviously a pretty significant problem. But that's not to say that LLMs are gonna disappear because again, as we start to build out new more deterministic models, they'll start to act as a reasoning layer in between the LMs. So I recently caught a blog that Eve, our founder actually hates this metaphor. I hope everyone here will bear with me. I use a lot of metaphors and I mix them. I'll do my best here. But I think the future of AI looks more like a sandwich. You'll be an engineer and you'll engage with an AI system through the interface, which will be LLMs. You'll have a deterministic reasoning system that will basically be the protein, which will ensure that the LLMs taking in information correctly is getting prompted correctly, that the AI understands the prompt. And it'll do most of the actual reasoning for the answering of that prompt at that reasoning layer. And then you'll have world data and world models that we're young is working on. They're almost at the condiments. It's providing the flavor, depending upon what different inputs you need to take in from the real world to get an accurate answer. And then you'll have that last layer of bread, which again is an LLM, which is given the output to the engineer. Like that's where we're heading to. And I think that's exciting. Right now we're eating this giant bread sandwich. And that's because, look, building LLMs and getting us to where we are now, almost a decade later, is a massive undertaking. And as we move forward here over the next five to 10 years, you're going to see that this sandwich is going to evolve become much more complex and able to be deployed into much more advanced areas of where we live and work. Yeah, I like that metaphor actually. I was kind of impressed by, and I know that we all take these benchmarks with a grain of salt. But I was pretty impressed with the benchmarks that came out recently from you guys. I mean, it's pretty, pretty high relative to what I'm seeing out there. And the verification piece of what you guys are doing so that it really becomes crystal clear what, whether this is right or wrong, seems like a big piece of the puzzle. And so can you talk a little bit about that? Yeah, so there's a series of benchmarks online. We've learned from talking with Jeralis and people in Washington that people don't understand these benchmarks. But so you have every year there's a mathematical, like the largest mathematical competition, which is called the Putnam competition, which is run. And it's the hardest math questions that you can possibly put out in students compete to do it. And there's a benchmark called the Putnam benchmark, which is 672 formerly-stated mathematical problems. Up until about a year and a half ago, LLMs were really bad at solving mathematical equations. Because again, LLMs are predicting tokens in a chain. We don't solve math problems that way. You don't need to talk out loud to solve math problems. As a matter of fact, it's probably going to be detrimental to your solution solving set. So I think when we first started, we created an agent, which we call ALF, which is actually built on chat, GBT, 5.2, 5.5. We tested on different models all the time when they come out. And we run ALF against these major benchmarks. In my first start, I think Apple held the Putnam Bench, which was around 70% accuracy at the time. And as of today, I think we're at 98% accuracy. So we've solved 660, that the 672. On Verisoft, we achieved a 98% success rating on Verina, which is another benchmark that looks at verified code gen. I'm sure we'll talk about that later. We hit a perfect 100% score. And on LiniVal, we're actually tied right now with bite dance, which we think is probably going to be one of our biggest competitors in the long term, as we think about the competitive landscape. Again, I'll show you a bit back to it. But these aren't vanity metrics. Formal benchmarks, they're really useful because they make failure visible. Hallucinate an answer in a chatbot can sound really impressive, but a failed formal proof. They don't get partial credit for confidence or for half answers. Either checks out or it doesn't. And that's why these benchmarks are so important for critical software development, because you can't have tiny cracks anywhere in the proof. It has to be 100% accurate because we're talking about systems where literally people's lives might eventually be on the line. And so competing and showing how the system is doing against these benchmarks is critical for us to ensure that we understand our formal reasoning as accurate as possible. And it's also helping us build out a massive database of theorems and lemas that we can then go and use to train our own in-house energy-based model. So a couple of things. I mean, I want to get into code and all the increase of activity in the last six months around code generation. But I think there's a lot of questions around what an energy-based model is. I mean, I know you guys use that term. Well, you just define that for the audience, because it's something that I don't think people really understand. Yes. And I am going to sit here and just pray that the next call have with YAH that I'll get hammered for being extraordinarily inaccurate and wrong on this. But when you think about energy-based models, I think most people think that to the old days of Boltzmann's machines. So let me start with LLMs. Could be all for the most part understand them. And LLMs is trained by taking large swaths of data traditionally from online. But now I think we've gone so far as we're now scanning books just to load in and creating synthetic data to now feed into LLMs. But in basically amassing the wealth of the entire world, which includes all the garbage in the internet, it becomes a really powerful, almost nuclear predictor of words in a chain, which helps at mimic intelligence, right? It makes it sound really good. But the problem is you get a lot of noise, right? Which is where the cracks kind of come in. Again, the more garbage you kind of train off of, the more flexible the systems can be and what they talk about and how they talk about. But the more potential for hallucinations also sort of start to erupt on. And we've created agent layer to try and solve that, but that comes a great cost of compute, right? Because in the end, LLMs are giant decision making trees. And the more complex decision making tree, the more compute power that we need. So energy-based reasoned miles of Boltzmann machines had a very sort of niche role in the beginning. You take data that you really want. So very specific data sets. And then you build it around constraints. So what exactly do I want to do with this data? And then you use a Boltzmann machine, which actually utilizes theories of physics. And it creates geographical landscapes. Outcomes that are, let's say, not preferred, like a, allowing a backdoor into a data protection system. That's very high energy. Yeah. Having a watertight system, which ensures that there's no backdoors, would be very low energy. And you prompt that EBM and it searches this massive landscape to try and find the lowest area, which means it's going to be the outcome that is most preferred. Jan was the sort of lead in taking like Boltzmann machines and really making them much more efficient because you create these giant energy landscapes and they're actually really compute heavy. But it allowed them to not have to like operate in a maze like an LLM, which thinks back to every maze it's ever run and said, OK, I have a thousand mazes I've run, I should go right here and see what happened. And if it bumps its head, it goes back to the beginning, it says, OK, I have to go left here. And it says, OK, I'll be sitting the thousand things. I need to go right again. And we'll just hit hitting its head and wall to wall to wall to wall to wall until eventually gets to something that looks like a solution set. Energy-based models have a different approach. Because you control, you build the landscape yourself with your own data, your own constraints, it's able to see the entire maze from the top down. And then it scores all the different pathways. And it says, this is the highest scored pathway to get you to your result. The other problem was, so you create this massive environment, which obviously keeps you heavy, but you can have multiple different lower sort of areas of the map. And so sometimes it goes to a mostly right, but wasn't it perfect. And Yalma is the one who really helped get us to a point where you can create a lowest point on the map, habit scored and sure it always gets the right outcome. But it also didn't have the ability back then to be able to take learnings from previous exercises and carry them on.
over. And this secret to what we call reasoning, what we would say is thinking. So our founder, Eve, a body, who has a PhD in tapal algebraic topology, and a deep background quantum physics, spent the last 15 years trying to understand how the brain works, trying to understand energy landscape, and also working in the quantum field, trying to figure out how you can take large systems and make them very small. And so over the last 15 years, she had begun to work on what we call an EBRM, which is an energy-based reasoning model, so that we can create an energy landscape that has a lowest point that engineers can put the data into, create constraints, and that the EBM will be able to navigate very quickly and very efficiently. And because it has latent space, it's actually learning, which means that when you give it something like a Sudoku test, which is a spatial reasoning test, it can solve it very quickly, very accurately, for pennies in a dollar, for what an Elon could do. And it learns from its own mistakes, which is something completely different what we've seen before. And so that's really the fundamental base on what we've been building, and why we're so excited about these systems. Super helpful to think about it in that way, and I think it's a good segue, too, to talk about how this plays out in software development, because previous to agents, software development was pretty deterministic. You run the software, and it either works or it doesn't work, and you can kind of tell. Now I'm seeing the scale by which people are writing code right now, is just mind-boggling. And recently, we're recording this late May of 2026. We've seen models that are being used to understand security issues inside organizations, inside code bases. And there's going to be a lot of work necessary to patch those code bases. I want you to talk about how these energy models play out in this landscape. >>Yeah. Okay. So, will you bear with me for a minute or two while I walk through this? >>Yes. >>Because I think it's important to help you understand our view of the world on this. So, most recently, the panel when you're listening to this, there was a major announcement from Anthropic around a new model called Mythos. And the key takeaway from Mythos is that this is for the first time a model that can look at code and can start to identify fundamental flaws in that code. And actually, this gets back to why joined Logical Intelligence in the first phase, what got me excited. When you think about how code is traditionally deployed, the vast majority of companies have a fairly simple process. You write the code and engine you write it. They test it. They debug it. They deploy it sometimes directly to the market, sometimes into a test environment. They debug again. You send it out into the world. And then you wait for a vulnerability to get exposed publicly and then you fix the code again. That's fine. For most systems, it's not great for the other like 49% of systems where you can't have vulnerabilities. That's what when we have those, you have the result of ransomware cases like one that I worked on, where MotionCourse couldn't brew beer for three months because of a little packed door and some old software code they had. That was then used for ransomware. You have pipelines being shut down. You can't have this stuff. It's one thing when you get angry because you're uniting app crashes at the last second because it's a little bug. It's another thing when you literally can't go out to have a nice cold beer with your friends or even worse, maybe can't turn the lights on. And so we feel that that environment is coming to an end and AI will empower this. So how do you fix it? Well, there's something called formal verification. This is a field of mathematics. It's actually existed for some time. There are some companies that utilize formal verification today. Companies like Amazon and Microsoft have in-house formal verification teams. So what is that? You actually have in-house mathematicians, PhDs, that take your code and break it down and they mathematically prove that the code does exactly as intended by running theorems against it. It takes a long time. It's very expensive, but it's why we don't tend to see too many massive hacks or breaches in code coming out of those organizations. The problem is very few people can actually go and engage in that because it's so expensive. And it takes a long time. And look, testing is great. But imagine if testing, if you're at a hotel, testing is going around and checking a couple of different doors to ensure the locks that we just installed work. Formal verification looks the locking mechanism themselves across the entire hotel and says does this locking mechanism hold up in all the environmental constraints that it might meet? That's a fundamentally different thing that we're talking about. And AI, because we're now able to start to solve some of these complex mathematical equations like we do in formal verification, is going to allow us to start to automate that entire process, which means that you will be able to go and lock down your code well in advance of having to deploy it. And as long as the prompts are input correctly by the engineers, we're going to know that there is a really, really high likelihood that that tech is particularly critical elements of that software is going to be iron, iron clad. So that's massive. So now let's think about it in terms of cloud code. Okay, code generation, everybody is using it right now. It's great. If you want to go and build like your own like website, you want to go and build an AI bot to comb through emails and float the ones that are important to the top like I do, it's really good. But the problem is anyone who's actually used cloud code and understands computer engineering, I'm not one of them for the record, but I hear it all day, my engineering constantly complaining about it. Cloud code and other code, an LM based code generators like it create way more complex code that is needed to do the specific tasks. Again, probabilistic systems, they create these kind of giant rats and massive code, which are kind of nightmares to have to go through and debug. Now, if you don't care if your personal AI assistant crashes, you don't really mind that too much. But as more and more engineers are going to use code generation software and they are going to use code generation software, it's trading this environment to where engineers no longer know the code that they're overseeing because they didn't write it themselves. And we talked a lot about how code generation was going to like and software jobs, you weren't going to need, you know, coders anymore. And this is like this new massive sweeping change in the economy. I think the reality is all it's done is take an energy from one place and move it to another, which is one of the greatest theories of physics that we have today. And you basically take an engineer's from like writing code to just going through and debugging it constantly. And so what we've seen when I'm hearing from my friends and you've advanced me your factory spaces, when engineers utilize, you know, things like cloud code to build these massive like complex code bases, they actually create more problems for them because when they do have to go deep bugs something, it takes a lot longer to go and find it. And you just are incentivizing maybe the wrong skill sets eventually, though, because now we're able to move into a more formalized environment created by AI companies like us are working on what we call formalized code gen because once you can, once you can automate formal verification, we can also automate formally verified code generation, which means that, you know, in the next like three to five years, you're going to have engineers are going to come in, they're going to tell their code generator, I want to build this, this, and this. And then I need a formally verified system for encryption. I need a formally verified system to ensure that HIPAA is protected and they're going to get formally verified code created on the spot so that we don't have to worry about code being deployed. So where does the EBM come in? Well, as most people know, compute costs are starting to go up. Oh, yeah. Again, it's me. It's me. Maybe I'll be wrong. But have you noticed that the prices are starting to rise in the last like a couple of weeks? Oh, yeah, absolutely. Maybe that's because both the two largest labs are about to IPO. And I do know a little something about public companies. I think for a long time, code generation has been a little bit of a loss leader. And that's okay. You want to get as many users you can in. But once they go public, having your core product be a loss leader or even just kind of breaking even like that's going to be tough for investors as well. Oh, yeah. So prices are going to go up. And if you think code generation today that's like at vast 60, 70% accurate, if you think that's expensive, wait to see how much verified code generation is. It's really expensive. Particularly on LLM. So for us, the EBM is all about scaling. How do we scale formal verification and very formally verified code gen? So I mentioned earlier that spatial reasoning is what the EBM does really well. So you have all these toy benchmarks that we take AI systems and we train against we test against one of the toy benchmarks that we did was Sudoku, which sounds really simple. But LLM is really bad at it. If you turn off their ability to just brute force it, they can't do it. And so we went and put out our demo to the world, which you can go to our website and see now that I said that I had to keep it up forever. It is what it is. And we just allowed people to go and create their own Sudoku or generate a randomly hard Sudoku and just watches our model or EBM, Kona plays the other leading LLMs in Sudoku to see. And after a week, we had around 15,000 people come and test it. And the end result was that we saw the little over 98% accurately Kona did. And it wasn't 100% only because some people go and create their own Sudoku, they weren't solvable. And the LLMs combined sold about two. The compute costs for us to run those 15,000 games for the LLMs was $14,000.
Boaz, do you want to guess what it cost us for Kona to solve 98% of those tests correctly for a week? A tenth is my guess. Four dollars. That's incredible. So for us, that's the future because if you can't be both deterministic and also much more efficient, we don't see how you can really take AI and push it into much more complex environments without creating this like Drake Honey in world from like an Orwell of just having data gigabyte factories just underneath every home across America. It just won't work. But this is what's exciting about AI. It's interesting that you say that because I'm hearing a lot of folks talk about orchestration layers being able to choose models that are less expensive, be really careful about not using the best model just for everything and this is really taking that to the next level. One of the things that I was surprised to hear you say is that bite dance that you think about them as a competitor. I want to hear you talk a little bit about that. But also, you speak to a lot of CTOs and when you look at the competitive landscape and you start talking about the work that you're doing, when do you see them lean in? Because these guys are listening to 10 pitches every single day. It's like when do you see them really leaning in to your message? As soon as you start talking about a termistic AI, so we released Kona the time of this filming about three months ago and the amount of outreach was overwhelming. We're still start up and the tech is still at its like early stage. We're now moving to what we call the design partnership stage where we start to because again, EBMs don't get trained by just downloading the internet. We actually need to work with engineers to get access to their data, which they continue to own, help us build constraints around the model and we test it to understand what it does really well. And the reason why we're able to go and get these design partnerships and actually start to test these model in some real world environments is because for the last five or six years, every CTO has been under pressure from their CEO, from their boards, from the market to have more and more and more AI across their systems. And the reality is that in most of the systems that matter, they can't have like probabilistic systems. Let's put LLM to sign. It's just a probability factor of this. And so as soon as we start talking about the termism and the fact that they get to control their data, the fact that they get to establish the landscape, place on what they believe is important, it's not one AI for all. You get your own model that you get to custom build, you get to go and work around and deploy as you deem fit. As soon as they hear that, the light bulb goes immediately out because this is what's been missing in the marketplace. Yeah, that makes perfect sense. Well, I always like to end the show with the same question. And that question is when you think about the future of AI, the future of work, describe it in two words and then you can elaborate on those two words. Chaotic determinism. Here's why I say that. I think it's pretty obvious that we're a big believer in deterministic AI because the market believes in it. The market knows that it needs it. The path to that is going to be uncertain for the next couple of years. I think that the market is really good at knowing when there's a gap that needs to be filled. I don't think they're great at knowing how to fill it right away. And so right now we're one of the few companies that are really engaged in this work. Looking at alternative architectures, again, Jan just launched his, just launched their company, Avi Labs, which is working on world models and other elements of alternative architecture. And really kind of it right now, except now we're seeing China is obviously like starting to speed up the development on this, which we expect is going to be our biggest competitor. And you would think that the light bulb would go off in Washington, DC, but I spent a lot of time working in Washington. And I was a G. I worked in the government affairs team and was actually lobbyist myself. It just takes Washington time to respond to this stuff. And that's by design, by the way, like we don't want government to just change as quickly as a startup does. But that means we're always a little bit behind the April. And when you have foreign governments like China that can just basically deploy all of their research towards a problem like alternative architectures and then could basically tell their labs and some of their state back companies that you're going to use this now. It gives them a tremendous leg up. But for us, this is a race to see who gets formal methods right first because the market advantage you're going to have is going to be astronomical. And the ability for alternative architectures like EBMs to be able to optimize advanced manufacturing in real time. Not these one, two, three year development arcs that we have today, I can test you for my time at GE, but to actually have an AI, a robotic system that is optimizing in real time for whatever task is working on, that's going to make things a lot cheaper for the economy. It's going to make goods safer and it's going to ensure that people are able to go and focus on the things that you should be doing. Everybody I think is terrified over what AI is going to mean for human capital. What is our value? But I think if we look at LLMs, LLMs are not replacing quants at finance firms. They're replacing like low level analysts. They're not replacing artists. Like sure in the movie industry, we saw AI Slap kind of have its, you know, it's heyday, but I think inherently human things are going to go back to being driven by humans. And things like really, really advanced manufacturing and machinery optimization, that's what's going to be driven by AI. And that's probably what should happen. It's just going to be a winding path to get there. But it's a better place for us to be than I thought that I think people thought we might be in a couple of years. So that's good. A little chaotic, but we'll get there. Well, it's such an interesting journey you guys are on and I know the audience wants to follow along. I was really impressed by your blog. You guys are putting out a lot of writing about this. Can you tell folks how to keep a track of what you're doing and how to get in touch with you or other folks in the organization? Yeah, look, logicalintelligence.com. We try and put out as much content as we can to help people understand some of these systems. You'll have a mix of really advanced technical writing from our CTO, from our founder E-Body and she's also on Twitter. Eve loves Olive and she is personally growing that Twitter account. And if you send her a note, she's going to respond to you. But you should always go there and check. We try and really keep a really healthy amount of content there that's both the average person who's like me and doesn't have a PhD in quantum physics can follow. But that also like researchers can engage with and help us think through like tougher problems that we try to solve. Love it. Well, Patrick, thank you so much for being on the shift. The AI podcast is really fascinating and it really helped me kind of better understand how you guys are thinking about this space. Thanks, Boaz. Really appreciate it. That's a wrap. Thanks for tuning in. It was such a pleasure to have Patrick Hillman as our guest on today's episode. Patrick brings a rare perspective of the blend's crisis management instincts with Frontier AI strategy. His insights into deterministic AI and formal verification and why critical systems need certainty rather than confidence are invaluable for any technology leader navigating this fast moving space. If you want to stay connected to Patrick, you can find them on LinkedIn. If you're building in the critical infrastructure space and want to understand how deterministic energy-based reasoning models differ from probabilistic LLMs, logical intelligence is worth a direct look. I continue to be amazed by the guests we've had on the show and I'm excited about the ones joining us in the near future. I truly appreciate you spending your time with us. Thank you for listening to this episode and don't forget to subscribe to Shift AI on Shift AI.fm, Apple podcasts, Spotify, YouTube, or wherever you listen and please rate the show with a five-star review. You can also check out our substack at shiftdi.substack.com. We're always adding new content so please take a look. Shift AI is syndicated by Geekware and our show's theme music was created by Dave Angel.
Podcast Summary
Key Points:
Patrick Killman's career spans crisis management, cybersecurity, and AI strategy, including roles at Edelman, General Electric, Binance, and now Logical Intelligence.
Logical Intelligence builds deterministic, energy-based reasoning models for critical systems (e.g., power grids, healthcare, manufacturing), where probabilistic LLMs are unreliable.
LLMs are probabilistic "guessing machines" that excel at communication but fail in high-stakes scenarios; even 99.9% accuracy is insufficient for tasks like surgery or autonomous driving.
The future AI ecosystem will be a "sandwich"
Logical Intelligence's benchmarks are impressive
Energy-based models (EBMs) use physics-inspired landscapes, where preferred outcomes are "low energy" states, enabling efficient, constraint-based reasoning without the compute-heavy trial-and-error of LLMs.
Formal benchmarks are crucial because they make failures visible—unlike hallucinations, a failed proof gets no partial credit, ensuring reliability for life-critical applications.
Summary:
Patrick Killman, Chief Strategy Officer and COO at Logical Intelligence, shares his unconventional journey from crisis management to AI, emphasizing the need for deterministic systems in critical infrastructure. 1% error rate is unacceptable. Logical Intelligence addresses this by building energy-based reasoning models (EBMs), which use physics-inspired energy landscapes to find optimal, low-energy solutions based on constraints, rather than navigating token-by-token like LLMs.
This approach is more efficient and verifiable, as demonstrated by their strong performance on formal benchmarks like Putnam (98%), Verisoft (98%), and Verina (100%). Killman envisions a future AI "sandwich" where LLMs serve as user interfaces, deterministic reasoning layers handle core logic, and world models provide contextual data, enabling AI to operate safely in critical systems. He stresses that formal verification is essential for making failures visible and ensuring trust, especially as AI expands into areas where lives depend on accuracy.
The company's team, including a Fields Medalist and veteran engineers, combines theoretical rigor with practical software development to push this vision forward, positioning Logical Intelligence as a key player in the next phase of AI evolution beyond probabilistic models.
FAQs
Logical Intelligence builds deterministic, energy-based reasoning models for critical systems like power grids, healthcare, and advanced manufacturing, where probabilistic LLMs are unsuitable due to their potential for errors.
LLMs predict tokens in a chain using vast data, while energy-based models use constraints and physics-based energy landscapes to find optimal, low-energy outcomes, making them more deterministic and reliable for critical applications.
Probabilistic AI, like LLMs, is essentially a guessing machine that can be wrong, which is unacceptable in systems like transportation or healthcare where errors could endanger lives, even if accuracy is high like 99.9%.
Logical Intelligence achieved 98% accuracy on the Putnam benchmark, 98% on Verisoft, 100% on Verina, and tied with ByteDance on LiniVal, showcasing their formal reasoning and verification capabilities.
Formal verification makes failure visible by requiring proofs to be 100% accurate, unlike chatbots that can hallucinate. This ensures reliability in critical software development where lives may be at stake.
The team includes Fields Medalist Michael Friedman, AI researcher Jan LeCun, and experienced engineers from companies like Google and Cruise, combining theoretical expertise with practical software development skills.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.