Go back

Data, Memory, Gyms and Models - NEURA's strategy

51m 45s

Data, Memory, Gyms and Models - NEURA's strategy

This podcast episode opens with a discussion on the difficulty of distinguishing real tech announcements from April Fools' jokes, exemplified by a satirical SAP post about fully autonomous AI factories. The conversation then shifts to the value of leaderboards for evaluating AI models, comparing the current LLM landscape to historical processor benchmarking. The core topic focuses on "Industrial Grade AI," emphasizing the critical need to reconcile advanced, probabilistic AI with the deterministic, local, and fail-safe control principles of traditional industrial systems like Programmable Logic Controllers (PLCs). The hosts highlight news of Agile Robots' strategic acquisition and partnership with Google DeepMind, speculating on the emergence of a standardized platform layer for robotics. The episode concludes by noting a humorous April Fools' post about AI-generated PLC code, which underscores genuine industry challenges with maintaining undocumented legacy control systems.

Transcription

8175 Words, 44128 Characters

English
This podcast is supported by Siemens, your partner for industrial grade AI. So have everybody and welcome to a new episode of our industrial AI podcast. My name is Robert Viva and it's a pleasure to talk to either a seabird. Good morning Robert. Long time no see no here. Good morning, good afternoon, good evening to all of you, dear listeners. Great to be back Peter. So we have a very interesting episode in the main part, but before we start with the main part, let's do a quick new spot. What do you have? I'm going to talk about April 1st, full day. Oh, okay. My latest hypothesis. I think it's going to be the end of April 1st, full day in a world full of hype and fake. How could we ever survive with or without April 1st? Not sure. Already a week or two back now, but did you experience one or two of those incidents? April 1st, full day, Robert? Normally my kids are involved in this, but this time they missed it. So it's quite a pain. That's almost already going to support my hypothesis, but then they are kids. I'm going to share one of you. So this is from SAP, a digital manufacturing. They say introducing SAP digital manufacturing, agentic AI addition. Now our AI agents will run your factory for you, full autonomous production, scheduling, self-reparing machines, robotic workforce management. And it's one of those AI-generated images. And you came over and you think, yeah, sure, why not? Everybody can say anything, right? But I didn't see anything weird about it until maybe at the bottom it says, "Come, April 1st." But I still hadn't thought of anything weird. What is wrong about robotic workforce management? What is what is put self-reparing machines? Until somebody was suggesting, and I think there was a, sorry, I didn't see the name if you are hearing this yourself. You were saying, "Okay, I'm going to laugh about this myself. We SAP, we were making fun of it." And I thought, "Are you really making fun of what other people are doing already?" That was a bit strange, I think. So bottom line still is, are we going to, is any day in the near future going to be an April 1st full-day? Because it's so difficult to skin through and recognize what is true from reality. Absolutely. That's interesting, yeah, because I read an article I think in the German hundreds blood that SAP is now working on a new AI strategy. Let's see what does that mean, yes? Yeah, you go, but it looks like at least one of them is not yet so fully convinced about their strategy they're making of April full-stay out of it. I didn't go into the detail. Maybe you can confirm that. I don't want to put wrong rumors, but there was a suggestion, there was an headline somewhere that their stock had gone down rather strongly. But again, I want to be very careful with this, because I did not check. I, every now, what I've been doing to be very honest is to check the NVIDIA stock for a year, every now and then, just because I know where it came from. That's even not specifically for personal reasons, for whatever. But it's almost for me that's like the, there was a number in my hat. And if I know and somebody is saying the market is going, you're not down or up or whatever, I just look at the NVIDIA and then I know, okay, if it's still around the number that I have on my brain, then everything is fine. If it's really below or above, and the last time it was totally above it again. And that's a number I've been following almost like for a year. So, were you reading the article? Was that correct? I think the most valuable company in the German stock market is now Z-Mense, I think. I saw that. Yeah. Yeah. SAP was the lead the last year, but they lost traction, I think. And for sure, when you see what's happening with Anthropic, we discussed this last time, they want to go into the tabular thing, right? Tabular data. That's SAP domain, right? So they need to act, I think. Oh, SAP is next to me. SAP is agreeing to the podcast. Okay. SAP. Hello. It's me, SAP. SAP for writer. Yeah. Cool podcast you're doing. Thanks. Thanks, SAP. So let's go on Peter. I was just writing down SAP's name. I was going to check with you anyway more off the podcast, but I was looking at the people coming to our AI and the Alps and SAP is, I wrote down SAP is number one. I always do it as you know. When there's new names that I people that I maybe do not know, I put a picture in accident. And SAP is number one without a picture. I know SAP from the very beginning. So great, great to see and hear that SAP is with you close, close to you and coming to us to the Alps. Yeah. I would like to briefly look back at our episode on leaderboards last week because I was surprised. I received a lot of feedback and a great deal of support for the researchers from Patabone. And I received an email and this is very interesting. It's written, "Hello Robert. I wasn't aware of this topic before and it's great that you also shed light on what might seem like niche topics." Because leaderboards are really niche topics, but I think it's a good way to describe our podcast, right? We also shed light on niche topics, not only agentic, but also niche topic when it comes to research, when it comes to industrial AI. And I'm really happy about this feedback because it's all about what we are doing every week. Thanks a lot. Yeah. Very good. So that is specifically the time series leaderboarder. Yeah, the TS arena, exactly. Because I do still look at what is the name of that other leaderboard. I've been following Peter. It's Faf Bench and Gif Eval. Yeah, right. And a little embryo. That's the one we're in beginning. Yeah. And I think there is a Peter. What's Peter? Peter Sautag? I don't know what is his name. Yeah. And Peter has been there and I think he joined him now. So that's where I don't really actively involve. I thought it was a great idea from the beginning where they did. And including now the time series leaderboard as well. Of course, they are very important. Yeah, but it's a niche topic. It's not like let's talk about new agentic workflow by I don't know who, but it's good that we also shed light on niche topics for our industry guys. Yeah. Most certainly. And I think I did suggest in the past that I'm being thrown back like whatever is up 30 years. No, maybe not as much, but maybe it is. And now there was the time of looking what processors do. There was a time I started with, 46 was just in the market. People who have heard about 46, there was just before Pentium. When I see the three numbers 48 and six, my brain brings me back immediately to that time right. And that was the time where everybody, I think it was like a Ziv Davis magazine. And everybody was talking about leaderboards on how to compare processors. And that time it was a time of Motorola. Motorola was the very big graphics. You know, there was no am I correct in saying there was no Nvidia, who want to be careful, but at least not under the name of Nvidia that was that was not. They were maybe somewhere on a very specific graphics leaderboard, but I mean it more in general. So how do you compare X86 and the new processors that were coming all the time? And that was that was very, very similar to what today is happening with our LLMS, I guess, right? But this was important for your industry, right? So you were part of this industry and you were you were looking at the leaderboards, but for the for the consumers over the customers was not important to leaderboards, right? In the end, not well, I mean, that's how we, I mean, we, I mean, I must say we at Intel at that time, we did this Intel inside. That was at that time for the consumer. We were telling the consumer, you don't, it doesn't matter whatever PC you're going to buy as long as there's an Intel inside. Now that's that's and of course a big branding thing. And I think at this time we are at exactly the same phase where the decision makers of industrial AI who cannot wait another year or they don't want to I mean they already started and want to continue doing things. They need to be able to have a good feeling for what are the different solutions giving me. And then at some point in time they're going to make their decision. One is going to say I want to have the best that has been the best and maybe and if the next week there's going to be another one the best I'm going to take that one. Somebody else has different reasons for choosing the one that is maybe number two or number three, but still very good or whatever. And in that perspective it is the same as the IT decision makers, maybe 25 years. ago and still today but today in a market that has very mature there's other reasons, there's other reasons why you're fixed already in a certain architecture maybe and I think with the market in which we're in with AI with LLM's with all the things taking place it's very very useful that you can go and look and see who's doing more. So what else do you have Peter? I for me the next topic is about getting deeper into the meaning of industrial and great AI, H AI. There's Stephen Yeats, he started a discussion, he says PLC's got one thing right that the H AI industry is unlearning which is about keeping authority local. Now you know that H AI has been one of my you know kind things for a very long time. I think we started with H AI. Yeah, thanks. So actually yes, meanwhile still at soft thing and where we're doing the industrial we started with H at that time. Now somewhere in the discussion there's a guy called Kiran Zebach, he summarizes the discussion nicely, he says AI can be probabilistic but the local trust and the enforcement boundary cannot. That was a beautiful set when I read it and before in my words and I want to share those I said you know assuming AI adds value right because if not there's no need for going to be blowing it right. Its execution still cannot be less than industrial grade. I use the term industrial grade that we always say Boris from Siemens has come up with right. EG by means of a deterministic government architecture. I want to step back a little bit. It's very interesting I believe to understand where Steven is coming from. Steven, he started his career as a hardware engineer and it's almost like getting back to what we were just talking about. Designing PLC CPUs at GE Fanook. There was another discussion that just comes to a brain right now. It's always about ITO2. This is a little bit separate. I come back but it's exactly the same topic and then I was looking back and because I said IT and OT have been separate since they were developed and then that was very very interesting. The first PLC was actually before we had a PC. It was the no oh I'm not sure which one it is. I think it was the Mart Bus. I'm not sure which one it was but it was 1968. We had already and they were not like on a microprocessor but they were on different subsets but it's very interesting to know that at the time that Intel was doing the I just said 46 and then it was a 4,000 for processor. Other people in a different part of this world were designing their PLC CPUs. In this case, Steven was one of those guys and he says the PLC sat and still sits in the cabinet next to the mode to develop the conveyor and it's running its letter logic. Locally, whether the network is up or not and that is the big point right. Even if later on there was a network and it was down, the PLC knew exactly what was happening. The authority never left the cabinet and now he says if we're going to go with H AI then we're going to assume that the connectivity between what is happening on the side of what has been the PLC and all of my elements in the network is going to sit somewhere else right and he says his point is the edge needs to become its own governments domain which in the past the PLC has always been. Now maybe it sounds a little bit complex maybe but I think it's a very very important topic, the idea being that if today something happens in your production environment right then the PLC knows about it, knows that something is wrong and makes a decision if with H AI something happens at the edge and your network is not connected to a central government then maybe the edge is still going to continue to work. That's I think that's the basic idea and the central government organization is not aware of it, who is not aware that something is happening. So I think it's a very important topic so I'll Steve and who is the CEO of Fedorant to come and share his vision in a podcast session about what he calls connected autonomy for H AI. Very interesting concept. A very important as well. Where did you find it on LinkedIn? Yeah, typically, typically those things do happen on LinkedIn. For me it's the main place where I got the information, yeah I'm not sure. I mean you'll go more to conferences. Conferences? You went to the robotics conferences I believe. Yeah exactly. I do that. I do that last and over we've been in the past or the I'm not going to be there this year but I did see for whatever reason I see a lot of a lot of what should I say messages with regards to you know coming to an over please join me but so I was just saying LinkedIn is the major source for what I see happening in the industrial AI space. It's very interesting the concept. I'm really looking forward to the episode with him because he should explain what does it really mean. Yeah exactly and as I said it's a guy coming from designing but at the same time you can expect questions from me to him and say you know what about the ladder logic was then you know invented at the same time as with the together with the first PLC right and that's where everybody now is saying and I'm very completely open to that and that's what I'm going to be asking him. What about new young people interested studying learning but at the same time you know the older generation like himself who now sees the capabilities of AI and saying are we going to stick to do we need to stick to the ladder logic and of course that's a huge discussion with you know 99% of the people or 100% of the people having learned to deal with the ladder logic and then now people saying well I don't need the ladder logic anymore but at least I'm saying we need to be in the end where the Robohydzer wrote I believe that is kind of the discussion of what we call industrial gate you know we need to be deterministic we cannot be probabilistic when it's say okay one time it works and the other time it is not we need to be able to so I think the potential of AI is to do all kind of things very specifically maybe around the the the generative so the the Jakub Tomchark side of the AI would say it can help us to marvel as great things but it I believe it has to take place in the end when we do the execution in a deterministic way with with uncertain boundaries what else do you have Peter I have two more or two more one is yeah one is an April one but it's not an April one fools message at least I believe not you tell me Munich based agile robots they close the acquisition of the T-Sync group automation engineering now that had already been announced end of 2025 not sure if we announced it at that time it wasn't in my brain specifically now what is interest was robotics yeah so it's the decent group automation engineering more than 75 years engineering including robot systems 650 engineers 10 global sites and agile robots is a seven year old DLR stands for Deutsche Luft und Raum für German Aerospace Center spin out two and a half thousand employees is that correct yeah the DLR yeah I couldn't believe that they are sprubic focus on advanced robotics to make a eye manufacturing smarter more flexible and they having a quality former a franca robotics right now following is acquisition the renamed now called Krause automation together with agile robots they will shape the next generation of physically eye and industrial production you know that's exactly what everybody's doing I don't mean that in a negative sense but the physically sense Jensen has brought it up when did he half a year ago was it last year? the number oh yeah yeah last year I think yeah but it's amazing but I really really mean that I mean if if you look at more you and I started seven eight years ago and now everybody everybody doing industrial AI and then the physical AI in industrial production right now the second thing related to agile robots I saw what was their strategic strategic partnership with Google deep mind and that's kind of what we talked about two weeks ago right because in our new section we put up to a high part of the Google intrinsic yeah we put up to hypothesis and I've seen it again somewhere else the discussion around it you know is Google going to become yet another robotics company and I said I don't believe it I believe they're going to become the Android for robotics platform layer and at that time that we were refer to this intrinsic to add some kind of activity up on their website. And this seems to be like an other announcement supporting. And then there were other people saying, oh yeah, anybody, there's so many companies who trying to put this platform layer there. And then I was asking, somebody else was suggesting, well, who are there potential contenders, AWS? Did somebody suggest Tesla as well? And then I thought, well, okay, well, if there's one company who is capable of really doing it, who's done it once with Android for the handheld, what is it? 90% of the worldwide market, my feeling was that they could do it. And I was just going to refer to this as a potential confirmation in the direction of, not sure how many partners they have in the meantime, that they could be doing this Android for robots layer. Not sure what you think of that idea. Yeah, I think I will talk to Lawrence. I think at the end of April, he was part of our AI and Amsterdam group. He's a R&D guy from Agile. And we will do an episode on that topic. And it's a perfect topic because in the main part, we have the Neurag Eyes also focusing on robotics. So, but we come later to the main part, but very interesting. I will ask Lawrence if we can talk a little bit also about the deep bind cooperation in the future. Yeah. Okay. Yeah. Yeah. Yeah. Yeah. Looking forward to who knows? So it doesn't have to mean that just because they did the handheld market they can do, but they have been in this in the robotics market quite some time. And so yeah, it's just it's just my feeling that that's what that's what it is. They are going to be trying to do the final one that I have here. Maybe before you then share what it is that you're talking about in the main section. There's James Booneyach. He launched a snap PLC. He says, just take a picture of your control cabinet and our AI automatically generates the entire PLC program. Very impressive. Advancet features reconstructs ladder logic. There it is again. From the way the wires feel emotionally connected. Listen, listen, what what the verveages write. It detects undocumented logic written by contractors in 2007. Wow. This feels like magic, right? It predicts which IO point the electrician meant to land. It gets better all the time. And it also it recreates the original programmers for process, including panic. I'm not going to say more. It's just you will record it. No, I will not because this was actually in April 1, 4th, you're okay. In case you hadn't learned yet. I think it was a wonderful way to to represent a frustration. And you can if you haven't been around engineers that are being called on site. All they find is exactly this ladder logic. And somebody touched the running system 10 years ago. And now for whatever reason, something has gone wrong. You do not know what is wrong. And then the programmer needs to go into. There is no documentation or nothing. But again, that's the thing I again will talk to Stephen about as well. It's finding out, is it? Is it do we need to go into this ladder? Which elements maybe just to round the app? Which of these elements are more now to turn at 180 degrees around? It's almost like these things are too good to be true. But some points of these I do believe will become true in one or two years. Really, this is meant to just make fun and joke. And the frustration of the electrician. I do believe that AI will help us. And I haven't been smoking. I do believe that I want to do this. Certain elements of this. Yes, exactly. That's what AI is going to be helping us. And I think the central question for this nevertheless is like, are we going to stick to the things that we have been doing? Or are we going to remove? Are we going to be doing things a different way? Maybe even. So with AI on top of it, let's let the AI take care of the ladder logic. That's a very, I understand that that's a very, very, what should I say, dangerous maybe. Saying because there's going to be now many, many, many of you saying, well, Peter, what are you talking about? You can't do that. But that has been said about, you know, coding years ago, you recall. And now, you know, everybody is kind of doing it one way or the other. So let's see if certain elements of this April's first full stay joke is going to, and nevertheless, going to become capable in the near future. Robert, what do you. It's a better joke than the SAP joke. Exactly. So let's move to the main part because I didn't interview with Jonas from Nora. And here's one thing I learned from the interview because he said everything we do in robotics will require now a very important new forms of memory and temporal dependency in the future. And this was very interesting. So new forms of memory and temporal dependency, very interesting episode. I think we didn't talk too much about Nora. We talked a little bit about Nora and their ideas, but we talked about models, world models. New forms of architectures, very, very interesting episode by Nora and by Jonas. Thank the lot Jonas. It was a pleasure. And if you or if he talks memory, does he mean like physical, hardware? Or did you talk about the more like logical structure of memory? Because that's the latter. If it's a latter, it brings me more into without knowing the details. But like, Alistair, I'm Xcelistian. I think that's different ways of dealing with memory as well, I believe. Exactly. From my personal perspective and when I talk to the robotics companies, we will see a revival, it comes to recurrent neural networks. So a new form, a new kind of RNNs maybe because the memory topic is very, very crucial when it comes to robotics in the future. Okay, looking forward to, I mean, you didn't talk about them, but maybe, but I'm sure you did asking about this company probably. Yeah, exactly. Because Neura, Neura kind of refers almost like to our brains. Exactly. And your neurological, I'm not sure if that is correct, but then they gave themselves their names. So it is maybe that's why he or you, he has been talking about the memory and the way that maybe, you're always going to be careful, not our human brain works because the people who know about had always say, well, still very, very different from the way that algorithms work, but okay, looking forward to, looking forward to listening. Peter, it was a pleasure. Thanks a lot. Robert, talk to you soon again and dear listeners. I hope you're going to be enjoying the main section and you already enjoyed our news. I have you with us soon again. Bye bye. My name is Robert Vibha and my guest today is Jonas Messner Jonas. Welcome to the podcast. Thanks for the invitation. Nice to meet you. So my name is Jonas Messner. I'm head of AI here at Neura Robotics. Yeah, I've been in AI pretty much my whole professional life. I've worked in automotive for five years bringing AI into various products that are today out there on the streets. And then yeah, here at Neura first took over the gyms and built the gyms, but then in the meantime, also took over the role of head of artificial intelligence overall. So essentially my responsibility is everything from data to training models to foundation models to deploy these on our robots. Okay, that's interesting. You are the right person to talk to. Today we don't want just to talk about robots, but more about your ecosystem, gym architecture models, foundation models. What are these gyms? Can you give us, explain us, what are these gyms? So essentially the gyms are addressing the big challenge of physical AI, which is scarcity of data. Essentially what the internet has been to large language models, like large amounts of text and images. This is not there for robotics or for physical AI because we're missing the real world interaction and like how we as humans interact in the real world, like the sense of touch or also audio or free division, all of that is not present in the internet. So essentially what the Neura gym is about, that we collect large amounts of such multimodal physical AI data such that we can even build real world physical AI. So this is what the Neura gym is all about. So you use your robots and the gym to produce data or M&Bron? Yes, so essentially in our gyms we have hundreds of robots acting in real world environments. So the gym oftentimes when we talk about gyms, people think, okay, this is simulation. No, it's actually real world physical infrastructure. So physical buildings teller operate our robots to collect data in real world tasks. So for example, you could think about logistics task where a robot is sorting packages or you could think about an industrial task where a robot is assembling parts and we're operating these robots in the gym infrastructure to then get the data out and train models on it. So it is a purely, let's say, data approach because scaling we learned that maybe the time of scaling is over when it comes to L&M, right? So are there a combination when it comes to simulation and using data or is it really only a data driven approach? So simulation of course also plays an essential role. But what we see for interacting in the physical world like for fine grant manipulation is that the real world data is simply irreplaceable. But we're having a multi-fold approach to our data strategy. So it's about simulation, it's about the gym data, so the real world physical data. It's also about a special concept that we have in the context of equipping humans with data suits and partnerships with data corporations. But it's like a multi-fold approach to data collection. Can you share a little bit about your approach when it comes to collecting data in factories or with who you are partnering to collect data? Yeah. So to collect data with partners, there are multiple partners that we're partnering with in a big way. So one of them or two of them that are public is Bosch and Sheffler. So essentially what we're doing in such partnerships is that on the one hand, these companies are of course using the gym. So they're using the gym to collect data on the real robot. But since that data collection part is not the solution to do everything, so there is more efficient ways of data collection. We're actually also equipping workers with a what we call data suit. So essentially the workers in factories they get a sensor on their body or sensor. So cameras, tactile gloves, audio and all of these things that also our robot has such that they can collect during their regular eight hour shift, can collect data that we can use for training. So this is a very efficient approach to isn't she get large amounts of data from real world factory environments. And the interesting thing about that is such data and doesn't exist anywhere on the internet. Of course, such companies are very careful about who they give that data to. And we're in the lucky position to have these strong partnerships with such companies and they essentially give us access to their factories and give us access to that data, which is a tremendous data source. Absolutely. But what do you do afterwards when you receive the data? Do you need the annotation of the data or what is the process then? So we're mostly working in this context with a self supervised approaches. So we're trying to reduce the amount of manual labeling by or to a minimum amount. So having been in larger cooperation I know that labeling cost can be a significant factor in such projects. And we're trying to use on the one hand self supervised approaches but also crown truth sensory around the humans but also around the robots in our gyms that essentially is overlooking the entire process. And that crown truth sensory can then be used together with also advanced AI models like large scale vision language models to auto annotate the data. So the amount of manual annotation that we do is limited to like really minor quality assurance and you know like the really tough cases where you still need human annotation but it's like a small amount that we manually label. So you mentioned you opened your gyms also for your partners or M&Rong? That's exactly the approach that we're taking. So while currently most other players are really focusing on solving everything all at once all by themselves which you can imagine I mean physical AI is huge right? There's thousands and thousands of tasks out there that have to be solved to make physical AI real. We're taking a very different approach to most other players here. We're essentially in a way crowdsourcing all of these data to learn the tasks. So the gym is really meant not for us to go in there and train everything all by ourselves but to bring in companies into the gyms. So we're bringing in huge companies from logistics, from industrial, from service, like from all kinds of areas and they go into the gym, collect the data for their tasks and then also via the gym infrastructure. So the gym is not just the physical infrastructure but also huge cloud infrastructure behind to do the whole data pipeline, smart training and deployment to enable these companies to train their tasks and with that actually do the whole scaling of physical AI. How do you convince them to share the data in your gym? So we see that for a fact everyone wants to go into the gym. So when we when we announce that concept actually we're we're drowning in requests to enter the gym. So because of these companies being all under lots of pressure as well to make their factories more efficient and with our gym infrastructure essentially providing them a place where they can train these robots. They are all excited about this concept that they want to actually enter the gym such that they can even train robots. I mean building something like the gym just every company by themselves doesn't make sense right. This is where the whole thing then again does not scale because if hundreds of companies all by themselves again builds such an infrastructure. It's very costly but if we built that once for everyone this is where the real scaling happens and where the real efficiency happens because everyone can use very similar pipelines, very similar training and data pipelines to train their tasks. You mentioned data pipelines training pipelines yes. We stop when it comes to annotation right. Can you go further in the process to explain a little bit what happens then? Yeah so essentially in the gym we have two kind of modes. So we're we're facing of course lots of companies that essentially have very little knowledge in the area of how such models are built and what we're trying to do here or what we do is that we completely abstract this process of data pipelines and training pipelines away from these companies such that they can fully do the operation by themselves. So they focus on the data collection piece and essentially then just feed that data into our gym infrastructure to train the models and all the magic happens in the back that's what we're focusing on. But to give you a little bit of details it's essentially I mean you you're getting that data you're annotating the data you're also in these pipelines you have curation steps that you can look into okay what's actually the relevant data that we should train on. Then there's the training pipelines in these infrastructure there's deployment pipelines together then either back on the robot itself or also depending on the use case and the task to deployment towards our cloud infrastructure being the nervous. Exactly that was my question. Is there a connection between the gym and the newer verse at the end? Absolutely I mean the neuro gym and the neuroverse are like one and the same it's like yin and yang one doesn't exist without the other. So essentially while the neuro gym is this data factory and the entire training pipelines to train these skills the neuroverse is the place where you in the end build your overall applications and you do fleet management and you oversee essentially what your fleet of robots is doing so they're essentially one and the same. Okay you mentioned now data you missed a little bit the whole topic when it comes to models right so which row do models play in the gym. Of course the models are a core piece. So it is always a model that's trained on the data right and here we're approaching that topic in in a multifold way. So of course we're using and we're building also on existing models that are out there from big players like Nvidia or also Google Gemini like making use of these models and tuning them to our needs that's one approach yeah but at the same time we also see that such models that are out there today mainly being vision language action models that they're functional in some cases like in the rather I would say easier cases but that in many more complex real world interactions these models face significant limitations because think about it us as humans do we rely on vision only to execute in the real world I mean no right We are really multimodal humans. We have a great sense of touch, which is usually a very big portion of our interaction in the real world. We have a great sense of hearing, plus also a 3D vision. And what we are building also by ourselves is own physical AI foundation models to address these shortcomings that also models that are today out there, considered state of the art that these have. So we are building, we are going even step further also on the model side and to solve many more complex problems than these VLA's, these vision language action models are able to solve today. But what is the problem with the VLA's? Is it a memory topic? Is it a sensing topic? What are the main obstacles? So it's pretty much the sensor modalities. So think about a task where you do not just pick up an object and place it in some other place, but you have to do some more fine grain manipulation in the real world. For example, you have to stick in a cable into a hole or something like that. There the force feedback or what you then feel on your sense of touch, on the tactile sense of your robot hands is essential for task success. Without this data source, you may still be successful in some cases, but your success rate in terms of task success is much lower than if you make use of this data. This is why we're in there are similar examples for bringing in audio where actually we have examples with some of our partners that show, you know, we have this welding topic where we see, well, the best welders in the world, they actually don't weld based on vision, but based on what they hear. It's actually fascinating to when we learned that because that's not what you would think intuitively, but it's actually the case. Absolutely. We had an episode on that topic, I think together with the rocket science guys in Switzerland, I think they're also using sound for welding applications. Yeah, I was also surprised to hear that the best welders on the earth can hear what is wrong. Yeah, and yeah, that's the three, you know, that's the, how us humans interact in the real world. We're a multimodal humans, you could say, and that's simply not covered in vision language action models. And you know, the nice thing about vision language action models is you can train them on internet data. And that's also one of the main reasons why you see all of the big tech companies doing that because they have access to these huge amounts of data like we as well do, but they don't have access to the actual multimodal data that you need to train such systems as I just explained them, but that's where we bring in the gyms, right? So that's how we are actually able to train such systems. But what about memory? I think when we talk about robots are the next generation of robots, we need more memory to understand what's happening in the environment. I did in the last two minutes or when I look back what I did five minutes ago, who I am, how I handled the situation. Isn't memory a topic for you? Absolutely. I mean, memory is a core piece in our own models as well. And that's actually also one of the shortcomings in addition to the multimodality of today's vision language action models. They're typically frame based. Now they have no sense of temporality and this is where we are also introducing essentially a sense of memory with the robot. Know what did it do two seconds ago or ten seconds ago or a minute ago? So absolutely that's a core topic as well. So coming back to gyms, is it an R&D lab for you, for your customers, building models and then publish models and then all of us is that the way to go? Yes. So it is a, yeah, you could call it a R&D lab. I would much more call it really a training world where you can train your use case, realize the use case, not just collect data and train, but then also validate that it's actually successful and then bring it into your actual application and your factory for example or in your logistics environment. And the way you do that is you train or you collect the data, you train the models, you validate them in the gym. And once you're happy with the results, you achieve like your 99% success rate, you can publish it via the nervous. And that's again another part where it becomes really magic is if we talk about, for example, a logistics partner that we have, they are training right now with us like package sorting as I described earlier. This skill of sorting packages is of course useful to many, many more companies. And they can then publish the skill via the nervous and even monetize the skill via the nervous. So it's also platform right to sell something. Exactly. You mentioned Nvidia and Google, right? So they have a lot of great VLAs, but you mentioned that they are trained on internet data, on public available data. Now you have this amount of data you are collecting this amount of data, you're building now as you mentioned, own models, those memory with more capabilities and you have this platform in the future or now and you already have it. Is it a second business model or is it a not to sell robot, but also to sell it say VLA models, new version, new approach when it comes to VLAs made by Nora. Is that the way to go? Yeah, I wouldn't say it's the second, but it's like directly in parallel. So of course we're selling our robots, like the full stack hardware, but our platform is, I mean already in parallel being sold, it's a second revenue stream, you're absolutely right about that. And it's usable not just for Nora's own robots, but in the same way also for other robots. And also here again, we do not believe that we solve every robot embodiment for the entire world, but that there can be other robotics providers that bring in their hardware, also bring it into the Noraverse, train their skills via the Nura gym, including our own foundation models that are just described and enable also these robots from third part. So for us, the platform, it's not dependent on being for Nora robots only. So it's an open platform, right? Absolutely. Yes. Okay. And the main topic is then that these models need to generalize, right? In different robotics application. How difficult is that? The way that this works is that most of the parts of the model are actually the same across embodiment, because it doesn't matter whether you have a wheeled robot or a legged robot or even just a single arm robot, as long as they have a similar camera or as long as they have a similar sensor for the touch or for audio, they will perceive the environment in the same way, the way they interact then. That's of course then slightly different between different embodiments, but the majority of the intelligence, the whole sensing and the thinking and how you want to solve things, that's the same across all embodiment. And in the end, it's rather like an, you may say, you may call it like an embodiment fine tuning that enables to do the right movement of the robot, but the perception stack is the same, yeah? We will see new, let's call them AI interdrators that are doing this fine tuning for the companies at the end as a business model. So I think the whole thing in the gym infrastructure is there will be companies acting in this in these gym infrastructure that may be an entire new market for companies to act in, to actually enable companies to train tasks or to collect data for certain tasks, to then train the models and then to bring these models on the robots, because it doesn't need to be of course the logistics company that goes into the gym, they may actually subcontract that to data and training companies or to service companies that do all of that in the back, but still they would do that in our gym infrastructure. And the gym infrastructure does it cost me? What is it? Is there a price tag on this gym infrastructure for me as a customer or as a partner or as a client? Yes, so there is of course a price tag behind, I cannot talk about the exact number now here, but what we do essentially is especially now in the beginning, if you buy a robot from Nura that comes with a certain amount of say training capacity and data collection capacity that you can do in the gym such that you can actually enable your robots to train and learn the tasks. Okay, Jonas, at the end, what is on your AI agenda? So you mentioned models, new kind of models, what else is on your AI agenda? So for us right now, it's really about the topic of scaling and scaling this gym infrastructure to enable the entire world. And we're actually having already like our gyms being currently built up in Germany, in China, also in the US and Japan plant. So currently our focus is really on scaling. On the AI side and in terms of technology, we're of course also following like several trends, like a current. And certainly many people are talking about world models that may be a completely new approach. And also there may be entirely new approaches to how to solve AI. The interesting thing here is that our whole efforts in terms of data collection and in terms of building the pipelines in the back, they will transfer one to one if then suddenly a new methodology rules physical AI. Like the data source that we're growing in the back that will transfer one to one. And that's also why we're focusing on that scaling so much. And not just worrying about, okay, what's the right model to right now succeed? Because I mean, this whole technology of foundation models is so young, right? It's a few years old. It will certainly change. The types of models that we see today will not be the types of models that will be the ones that solve things in three or in five years. But what will be the same is the data. And this is why we're focusing on that data part so much. And getting that right and getting the getting good high quality data and also the right amounts of such data such that even if the technology advances significantly, we're always having that huge source of data in the back. Okay, Johannes, I keep my fingers crossed for you, for your company, for your gyms, for your nervous, for your new world models, for your new VL8. Thanks a lot. It was a pleasure. Thank you so much.

Podcast Summary

Key Points:

  1. The hosts discuss the blurring line between reality and April Fools' jokes in the tech/AI industry, using an SAP AI announcement as an example.
  2. They explore the importance of leaderboards for comparing AI models and industrial technologies, drawing a parallel to historical processor comparisons.
  3. A significant discussion centers on "Industrial Grade AI" and the challenge of integrating probabilistic AI with the deterministic, local control required in industrial settings (like PLCs) for safety and reliability.
  4. News highlights include Agile Robots' acquisition of the T-Sync Group and its partnership with Google DeepMind, suggesting a move towards a platform layer for robotics.
  5. An April Fools' joke about an AI that generates PLC code from a cabinet photo is noted as humorously reflecting real engineering frustrations with undocumented legacy systems.

Summary:

This podcast episode opens with a discussion on the difficulty of distinguishing real tech announcements from April Fools' jokes, exemplified by a satirical SAP post about fully autonomous AI factories. The conversation then shifts to the value of leaderboards for evaluating AI models, comparing the current LLM landscape to historical processor benchmarking. The core topic focuses on "Industrial Grade AI," emphasizing the critical need to reconcile advanced, probabilistic AI with the deterministic, local, and fail-safe control principles of traditional industrial systems like Programmable Logic Controllers (PLCs).

The hosts highlight news of Agile Robots' strategic acquisition and partnership with Google DeepMind, speculating on the emergence of a standardized platform layer for robotics. The episode concludes by noting a humorous April Fools' post about AI-generated PLC code, which underscores genuine industry challenges with maintaining undocumented legacy control systems.

FAQs

April Fools' Day highlights the challenge of distinguishing between real AI advancements and humorous or fake announcements, as seen with SAP's AI agent prank.

Leaderboards help decision-makers compare AI solutions, similar to historical processor comparisons, by providing benchmarks for performance in niche areas like time series analysis.

PLCs provide local, deterministic control in industrial settings, ensuring operations continue even if network connectivity fails, which is crucial for reliable edge AI governance.

Connected autonomy refers to edge systems operating independently with local authority while staying integrated with central governance, balancing AI capabilities with industrial reliability.

Agile Robots, through acquisitions and partnerships like with Google DeepMind, aims to advance physically intelligent robotics for smarter, more flexible industrial production.

AI tools could automate PLC programming by analyzing control cabinets and reconstructing ladder logic, addressing challenges like undocumented code from past contractors.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.