Go back

EP19: AI in Finance and Symbolic AI with Atlas Wang

70m 34s

EP19: AI in Finance and Symbolic AI with Atlas Wang

In this episode of the Information Botanical Podcast, hosts Robert and Alan welcome Atlas, a faculty member at UT Austin and research director at XTX, to discuss NeurIPS 2024 and broader AI research themes. Atlas describes the conference as a dual role—presenting academic papers and staffing an XTX booth—which made it busy but rewarding, especially for interactions at the intersection of AI and finance. Alan shares mixed feelings: he enjoyed the LLM sampling research and San Diego’s atmosphere, but was frustrated by the poorly received conference app and the high proportion of non-researchers, including VCs, which diluted scientific discourse. Atlas defends VCs as technically savvy, noting one who had read his papers, and emphasizes the value of workshops for pure science, citing examples like the pluralism workshop with Ted Chang. The conversation shifts to Atlas’s research on low-dimensionality, a theme spanning his work from compressive sensing to pruning, low-rank methods, and symbolic learning. He highlights a recent theory paper on how neural networks can learn symbolic equations via gradient descent, arguing that the ultimate compression is converting neural knowledge into symbolic, human-readable form. Alan connects this to unsupervised learning and human cognition, though Atlas cautions against over-extrapolating biological priors. The discussion underscores ongoing tensions between commercialization and scientific integrity in AI conferences, while celebrating the enduring value of low-dimensional representations and symbolic reasoning.

Transcription

11361 Words, 62787 Characters

English
[Music] Hey everyone, and welcome back to the Information Botanical Podcast. And today we have a very special guest, Atlas, he's a faculty at UT Austin and a research director of XTX and he's living in New York City and also he's a good friend. Hey Atlas. Hi, hi, Robert, hi, I'm a thankful having me. And well, by calling me a special guest, I don't know which of the point you are feeling the most special. Is that I'm a faculty of UT or a research director of XTX or you're a good friend or a living university? All of them, but mostly that you're a good friend. Okay, then I'm honored. Thanks for calling me that way. Thank you. Hey, Alan, how are you? Yeah, and I'm great and it's great to be here in New York, also on the same city as you two. So we're all shivering together in this extremely cold city. And it's really nice to get to see Atlas again because we spoke briefly at NERIPS and now we get to kind of synthesize our opinions about that conference and share our minds. So happy to be here. It certainly feels like a sharp contrast to just coming back from San Diego to New York, a weather-wide. Yeah. So tell us, how was NERIPS? I'm so sorry that I missed it, but tell us. Did you like it? Did you enjoy? Maybe I can get started. Well, lots of people busy and I'm there within both hard. As an academic, I have papers with my student to present. For my industry affiliation, I work with the act yet, and we are a plant-home sponsor of the conference. So with all half a booth, I stand there for two days and meeting all kinds of exciting people. So yeah, that dual hearts will feel a lot. A lot of work. I didn't have even a chance to see Alan's side saying in the end of the year ago, but I'm still, but still, I'm blowing away by the warm weather and also the excitement from the people we talk to, especially on the interaction of General De Réa and a finance. I think that's definitely as full-light. I hope so in this year. I'll do so, Alan. Several things. One, I was very happy to see a lot of LLM sampling research, which is like a personal interest of my own, especially met a few people who I even wanted to bring on as interns, because thought works are companies pretty openly, you know, trying to create sampling lab, get all the sampling talent, gobbled up, because we think that it works well for consulting, among other things. And so I was happy in that sense. I think I was very frustrated that they didn't use HOOVA, which is the conference web app. The current one that they have was extremely poorly received and did not have it even close to the features. So I felt disconnected from the conference in a way that I hadn't before. And I mean, it was good. It meant that I had to constantly do in-person interaction to ask for directions or check billboards, like electronic ones. But it definitely felt like a step backwards in quality compared to NURP's 2024. That said, San Diego is a stunning place to hold a conference. The people there were relaxed in a way that I didn't expect for Californians. And the workshops were very nice. And yeah, the huge quantity of people. I do think it's a good thing to have more people, right, as long as we're not totally debasing the quality of publications. And in general, I think that problems with quality are not due to just like, oh, too many people are there. So I'm happy that it's big. Though I'll just say, I do find that NURP's attract way more not researchers in the other conferences. There's lots more money for good and for bad. I don't think of Elio Vaz, who would hate the people coming to here, coming to us and vice-souths who fancy dinners. I never hated that. No, no, I like it. But it also definitely means that a lot of like, if you just go to a random person there, you're chance of that person being a researcher with a paper is a lot lower than even at ICLR. And sometimes, sometimes I do want to eat with the researchers instead of the VCs. But I really, for somebody in your position, I think the VCs are even bigger deal. Well, first, I really, I mean, we're only swimming as into this full cast that I think I already see, I'm hearing Alan being sarcastic about multiple things. Which I was like, no, no, I mean it, I mean it, to be clear. I don't mean it in a bad way either. No, no, I mean, money. Yeah, good post cast should post cast that should do, right? I mean, above the meeting, I have had dinner with at least the two of the past PC channels and we all talk about the same thing. I mean, you certainly know that a lot in complaining about that. But there are other reasons, which I mean, I prefer not to discuss. And the California's work style, okay? I'm glad you're saying that you have been thinking so, and you got to be very different from a kind of spontaneous, euro work style, which I haven't, I happen to have the same stereotype as well, which is why I don't live in California. And the end of the last thing, the VCs. Okay, so first, the disclaimer, I work with VCs and XTX, the firm itself has a VC team and I love those people, great people. And I think those days, VCs come to conferences for, which are probably more technical prepared than some of my students, if I have to be honest. On my flight from New York to San Diego, I was sitting next to VC and he looked at me and he says, "Here, I have read your papers before." And I saw that that was a compliment. I said, "Okay, so what did I read?" And he said, "I think you write these papers, these papers, these papers, these here." I was like, "Okay, my student probably don't even remember I wrote those papers. Great, I'm honored." And he said he has to read papers every day and go to our cabs, go to Google scholars because the VCs try to propel them to be more discriminative about the good case and bad case, which is just challenging, just to think about how difficult we have to select good paper from bad paper. And now we're talking about selecting good money versus bad money. So I respect their work and I love that. A broader audience, I'm reading my and our papers. I'm probably less sarcastic on that point. Well, in particular, we know that a bunch of AI researchers, based on what happened with ICLR, are full on just AI-generated papers with a generated hallucinated citations, and similar lasings. And the reviewer is getting revealed in people trying to pay them off or whatever. We have a lot of terrible behavior and poor, shoddy science from our communities. I'm not surprised that analysts and others who need to make money are in some cases doing parts of or even our whole job better than we do. Yeah, but I think it's a good question. It's a good question what you actually want from a conference. I think, Tim, now, we have these research conferences that you only publish a work to other peers that you want to exchange ideas and make better ideas and understand what other people are working on. And we have these, I don't know, business conferences that someone wants to sell you something. And here, it's kind of like mix. Most of the time, you also want to sell your ideas and to exchange ideas. And you also, most of the companies want to treat it as recruitment event. And there are all these vcs that it's not believe in what they are doing there. They want you to stay up to date with the latest ideas and papers. And they also want to get no other people. And it's not clear what the goal is. And I think we have these mix and we need to understand how to do it. And what we actually want from these conferences. We as a community. That's right. I think one response, one personal reaction I had for this is first, I enjoy those big conferences are one trip, close the effective way to meet all my friends. Like I haven't been traveling to conferences in past one and a half years due to a new addition to my family. So this one has been a part of a renewed reunion of friends every day. And I truly appreciate this opportunity. Regarding academic, I mean, academic intake from reading papers. Let's be honest, most of us already know what new, this new episode is going to go on in paper wise before we ever make this trip. I probably six months before the conference, we already know what it would be the spotlight. It would be the spot with on papers, at least in our area. I, I, compared to the main conference post the session and post the session or others, I probably appreciate workshops even more. The main conferences because I feel, yeah, for conferences, it, for the main conference, it's always associated with many more incentives beyond the science itself, like Alan mentioned, there are people who try to press ulcers and reviews for bad reasons. They probably don't have the same bad reason for workshop, that makes, that makes the New York's workshop a bit more pure science. And people are also more brave in publishing their half-baked, uncompleted, frontier ideas there. That is, I appreciate the match more. So I actually spend a lot of time in workshops, more than I do in the main conferences. I find that, I find that is a balance that I try to strike myself with. But I don't know how, how, how, how. Well, I want to quickly note that the last year's Neurips had a pluralism and creativity workshop that Yamlik was doing. or sorry, Yashua Benjiyo and his students organized in ran. And beyond just having Yashua who actually stayed the whole time, which packed the room just by himself, they also got Ted Chang, the author of the book behind the movie Arrival, which was a big deal movie and Arrival's sequential fantasy for all, you know, pre-chat GPT AI researchers, right? Of, you know, decoding alien language for humanity. That workshop packed packed packed like seven or 800 people in a row. I'm pretty sure the fire department would not have been happy about that, right? How how packed that room was. But then he, you know, even at this NERIFS or at last years NERIFS, you go to other workshops that were not organized with the same kind of heavy hitters. And the room might not be packed at all. In fact, this probabilistic sampling workshop where some of the work that we did at ThoughtWorks came continued basically PELISD coding, which is like a subsequent method compared to minP sampling. You know, 50 people in that room maximum. So even among workshops, it's very feast or famine, at least in terms of who you can get your work in front of. And it's not even always clear when you're submitting what kind of workshop it's going to be. Like is it going to, I mean, obviously, if you see Yashua organizing, it's going to be big. But there are other workshops that I found that get large or get a huge and it's not clearly obvious to the people submitting that. Yeah, but I'm not sure why the workshop turned out to all the population sizes, a good indicator of the return you can get from the workshop, right? Like you say, there's a stochastic sampling workshop, or other more mathematical principle, foundations workshop. Sure, you will never get a big big room, but you get good audiences. This is always a signal to noise ratio, right? I was cognizing, I'm also cognizing workshops this year, gone generative, I have a finance and that's a very packed room as well. But I myself also emerged in a few other workshops. One is OPTML, Timulation for Machine Learning. I'm always in this workshop, since it's beginning. And I think I have talked more and learned more to this workshop than many others. We all have our subcommunities, and if I'm in a room of only 15 people, but those are 15 people who can 100% comprehend my work, this is the best workshop I want to be. Yeah, I agree. Well, talking about conference dynamics is an interesting thing, but I think you have a lot of, I mean, you're quite yourself and most of the students you work with are prolific and productive, and you have a lot of research that we'd love to start asking you questions about. So you want to be revealed to? Yes, that is. Yeah. Well, so I guess the first question is just like, what's your favorite thing that you've worked on, or even your students, if something your students have done is more exciting to you, I guess maybe this year, since it's nearly over. First, I like the way you raise the question that I try to differentiate my and my students research, and I think this is very fair. That's the question I like to ask my students, some of them are not faculty members as well. For myself, I always like one research theme called low-damageinality. I did my PhD in the background of single processing and optimization. And my first research topic was called Compressive Scenning. I don't know how many younger audiences are still aware of this direction, but that's still very deep close to my heart, fast optimization. And then they went to low rank and non-linear manifold, and I just keep enjoying doing that until after a few years, I realized I cannot publish a single paper because everybody entered a new area called Deep New and Network. So let's get to that new area too. But to my profound, to my profound joy, after I went to the one to exploring those overfound, overparameterized, the newer network, I find there are more opportunities for me to practice and actually deepen my understanding of low-damageinality, prominent examples, and may include pruning, lottery ticket, hypothesis, low rank, and mix of expert, and I have had great joy studying all those things, and by practicing what I have learned, the sparsity, low rank, and low-damageinality manifold. And this has been the live work that I continue engaged in my own interest and effort. In the recent years, my group, my student and I worked together on a few efficient inference and training algorithm, a large foundation model, some of them, I think a reasonable welcome by the open source community. But if I have to name one favorite work, I myself did in the past one year. That's probably a pure theory paper that haven't found much of the practical usage yet, but I like it anyway. The paper is called a paper analyzer. How a newer network could probably learn symbolic equations by gradient descent? So there are many recent deep learning theory papers study how gradient dynamics are going to drive a newer network towards the solutions of structural properties, like sparsity, low rank, low entropy, those are pretty known to the community. However, how to close the gap between continuous dynamics and discrete structural learning is highly obvious because we're always using tools from the continuous. Let's add less. Yes. Wait, let's go step back. What does it mean actually like symbolic neural network? Like how you actually look? Why do we need to care about it? Well, you want a logic and a rule, right? Logics, rules, and association relationships. Those are things that you could write with symbols. And I want to know why newer network is not just learning a huge bunch of black box functions. Instead, it's something that could be mapped to a low-dimensional symbol space. The reason I like that is because I consider this to be, I mean, I would say that without proof, the ultimate form of low-dimensionality. We use deep neural network to learn big functions, and then discover that those big functions can nail down to some much lower-dimensional scenes. Lottery tickets is one prominent example. And we have a lot of spars low rank, as a low-dimensional stuff, we can apply to deep neural networks. But for pruning algorithm, you can probably compress by 50%. For low-rank algorithm, if you are doing fine tuning, of course, you can compress by maybe hundreds of times, compared to using a formatrix. But what I would argue that the biggest compression of neural network is to compress a neural network into a non-neural network. That means compress what you have learned into knowledge that you can write with symbols, like what you have read on textbook. I would argue that human textbook, human knowledge that can be written and recited like we are doing now are the best form of a compression. And that will not be on your network. So that's the goal I said for myself. I don't want to compress neural network into another neural network. Distillation, low-rank, sparsely, that told me a lot of things, but now I want to compress them into a discreet form, a symbolic form, and eventually maybe even human-readable language. Again. So when you start talking about association symbols, I think association rule mining. I actually started my career in machine learning, at least with like my college classes on learning all the ins and outs of a priori algorithm and alternatives to it for better association rule mining. And I was always fascinated with that algorithm in particular how it relates to databases. It felt like automatic database constructs. Yeah. Yeah. And that I feel like it's not. We'll still use the else. We should get to memory right every day. Yeah. And beyond that, there's all these apocryphal stories about like people were target or whatever discovered that somebody was pregnant before they did and was like sending out ads for diapers or whatever using association rule mining techniques. I also hear you talking about dimension or at least low dimensionality. And I think dimensionality reduction, right? You mentioned like manifold methods. I am a strong believer in the importance of unsupervised learning and think that it's been unfairly devalued by neural nets, both dimensionality reduction and clustering on low dimensional representations. And I point out that models that perform basically the task that unsupervised is often trying to do at least an image space, the Sam model, right? Segment anything model. As a great example where the ideal, I think, image clustering algorithm does the same thing of segmenting your images in a logical way. And when I think about how humans learn, I claim at least like the majority of like a baby's early experiences are unsupervised, right? They have unlabeled data that they learn the relations to by just figuring out what objects are like this apple is not the glass on my table. Like that's all you need to be able to learn a lot of things. Because this is a life experience. And now I have every day with always my one year old. Well, I have to point through that apple and I keep telling my song. This is apple, apple, apple. I don't know how he's learning, but I'm trying to give him sound supervision. Okay, going back. But your label, but your labels are very sparse compared to the data that they do in an information theoretic sense. But it's not. Yes. Again, I also a big believer in set of supervised and unsupervised and the mention of action and all these. But like we need to make sure that we are not making this connection. and we are not taking it too far away because you don't know the power knowledge here, right? Like you don't know what is the power knowledge that humans have, right? Like you don't know how much, like again, prune formation encoded in the brain that you only need like some small amount of health information. But we have it as a whatever. Yeah. We have more of Xperia docs. We know that we've evolved to be able to do things like walk and run over millions of years, which is probably why it's so hard. One of the reasons why it's hard for AI. So we have echoes of evidence of what's what's prior and what isn't. What abilities probably prior. We all know that every cellary is wrong. Just a song I use for right? I think the same can be applied to mathematical priors. I was just talking about my journey over low-dimensionality priors, sparsity, low rank, manifold, and now symbolic. I mean, all I'm trying to describe this thing from different mathematical perspectives. But I don't think any of those is buried in our brain, right? Of course, there are biological sparsity. But really, I don't think when I was learning things, there is a L zero norm in my mind. Yeah. Yeah. But I see. I see. I'm excited by data. I agree with you. Sorry, rabbit. No. Again, I think that like, in principle, that's true, but it's not clear to me. The optimal solution that we may want, and there is the practical solution that involves the optimization process and the representation in how we define the thing, how the specific types of data and all these things. And I think we need, again, we need to be careful to say, oh, yeah, humans are like that. So let's try to push to this direction. But again, I think this is a good part. I'm like dimension reduction and compression and unsupervised and set of supervised learning is the way to go. But I don't think it's because humans are doing that. Well, I remember, yeah, I had the same as saying, probably more than 20 years ago, right? Like human-bured L plan, but it does not swing like part. It's a, yeah, probably many years ago. Yeah. Yeah. I agree. I do think as long as the information series, as long as the information series still holds true, low-dimensional, liquid-perseway to where we're there. And that's why I like your work a lot too. Yeah, I love it. Thanks. And no, I'm a big fan of compression, right? Like most of, like, a lot of my work is about compression. Okay, don't get me wrong. Okay, so we understand like why we want symbolic, quite to learn symbolic concepts and models tell us how you have what you, what you want to do with that. What you want to do practically or what you don't want to do prove wise. No, both. Let's start with practically. Then you are, you are effectively muted me because I don't have a practical application with us here yet. Okay, I'm kidding. Let me ask you. So let's do the theory. Okay, so let me, let me answer. Okay, let me still answer this from practical to theory way. So practice, like I said, I was initially driven by the desire to convert newer network to symbolic and two most obvious benefit number one efficiency because running let's say light GBM tree is being in anyway much, much easier than running any newer network. However you compress it over the steel that right? In general sense, of course, I'm not talking about specialized hardware. And in some extremely resource constrained latency sensitive applications. For example, CPU based I like in networking or congestion control those basic CPU tests. The symbolic equations are much better controllers, compelled to newer networks or reinforcement learning in terms of resource. So we did, we did some one of our early work was in collaboration with the ML6 folks are trying to first learn a reinforcement learning to control the networking congestion in CPU environment. And then convert that reinforcement learning into a decision tree. And if you care about speed, that's immediately 400 to 500 times acceleration on CPU environment. That's pretty straightforward. The conversion is done by Gaussen fitting basic, but basic way in symbolic regression. But we see that works. And that's the, and the Zinzata specific case is a better way better than adding network compression algorithm. We have tried it's not 10 times 20 times, it's hundreds of times. And then we and then we pilot the data to interpretability. We trend the visual reinforcement learning algorithm on opening a gym environment to let it play in the simple visual games. And then try to convert that scene based reinforcement learner into again, a symbolic tree based the pattern. So we see how this learned symbolic algorithm started by grounding color patches into object in a unsupervised way, actually, I would say in the end and and then use learned associations and logical operators to combine those to combine those basic units and make decisions. It's also a decision tree light thing and started from some like object partition, you can consider it automatically learns how to patch and patch and segment the input image patches into semantically reasonable subject and maybe that's maybe what we want to do in the middle of vision anyway. And then the decision making operator algorithm operate on terms of the small patches. And we have a paper on palm on that. Although I wouldn't say this, I have to be fail. I don't think this would scale up very well in complicated vision control based like we are just playing this in like an open mind craft type of simple visual images. But I think that again shows the shows the possibility that neural network are actually just learning composable logic plus perception, which plus the perception modular, which can also be described. And one way to do it now, appear to be the same or the same model like Alan just mentioned, you can we can convert the continuous things into discrete things like we're doing with grounding in robotics. So we are and yeah, we have a bit for a few more work along this slide on different applications. Yeah, I think efficiency and it practically has been two main driving forces. And we are not allowing doing this. We have seen great work from flat flat iron Shelley host group has been doing a lot on that and and the microscrupter in Cambridge has been doing the symbolic of regression. We we like that work a lot too. Seeing all those sero in empirical work, I have been holding this question from yeah, from I think from three years ago we have been discussing with my student quite extensively. Why is this possible? I mean, we first sure neural network is a universal estimator. If you try to be lazy, you can use this to answer Alan's in deep neural network. But the true question is why neural network would kill to learn Alan symbolic formations. For example, I reduced the order polynomial. Why it has to be this way? I always want to use a simple example, not fully precise, but a simple example I used to ask myself when I was a high schooler. Why Newton's second law is one over r square? It's not one over r square 2.015 something. Why this is a beautiful simple R square? And if physics PhD could have done integration to show the R square must be there. But if Newton was using neural network to fitting this, actually I'm talking about a capitalist word. But if Newton were to use today's neural network to fit this equation, the neural network probably will not tell Newton this is one over r square 2. It will appear some well the equation because neural network does not build with the sense of learning clean compact equations. And even this notion alone is very hard to describe how you should tell your new network to fit a function that is clean class function. What does it even mean? I think there's a lot of human prior here. So I was trying to first try to solve my student I'm trying to start from analyzing simple theoretical cases learning a polynomial from since for learning a polynomial from synthetic data. And we're trying to learn the perblons whose underlying structures include compositional algebraic and logical relationship between input on the around beyond just doing statistical recognition. And we have been we have done some simple theory assuming some basic mathematical mathematical group operations and ring structures and proofs on theorem to show that actually if your data generation process follow algebraic structure, then we can prove neural network. Indeed is able to discover that underlying symbolic structure precisely by using gradient descent, which was amazing to me because gradient descent is a continuous dynamics and you're trying to reach a discreet target. And we show that they can converge and that is done by some mathematical tools. For example, measure space analysis was instant gradient flow of functionals. I would be happy to throw you the link but probably let's do not spend the whole time on the proof here. And sorry for getting me a bit more excited about that. Sorry about that. No, that's great. But like so, okay. So like we have these like you proved right? And you also showed like something that later that like you can have discreet problem and with gradient descent that is like you can actually converge to the real solution right to the optimal solution of this. Yes, we show they can find a solution but the cabinet, I mean beyond all other things you would expect from our first I think of first of its kind of proof is very preliminary with a lot of assumptions. I have to confess on that. A major caveat here is our theory does not prove, does not provide a construction. It is a existence proof. I'm not saying that I can get rid of the neural network and directly reach that symbolic equation. I feel like the situation in low-tritic is hypothesis. I know there's a lot of tickets that exist somewhere in the initialization, but there's no good way for me to pick it up. So I think there are way more works to do for us to both push this serum into a really practical verify bookcase and to directly reach the symbolic equation without going through over-primed arises training. I think later would be a Holy Grail problem. I want this looking to next. Okay, so and I want to ask you, what next? This is something that is generalizable to other problems. I think this is something that, this is the way to go. This is the way to go to try to formalize concrete problems in very regards way and to show that HD can convert to these problems. I don't like to think this is something that we need to go to other way. I like to try to to make relax these assumptions and maybe we will not show it like rigorously in a mathematical way, but it's still fine and like we have some approximations. I think I already, I think I actually already answered that in how I approach this symbolic purple, basically do those things in parallel. We do those empirical proof of concept and we do the theory in parallel. They one does not necessarily wait for the other. I like doing rigorous theory, but I really haven't seen like, I mean, let's put it nicely. I mean, it's very real that how rigorous the theory will drive our empirical advanced those data. So I prefer to do them in. So what kind of problems do you think you actually show it? What type of problems, even like if you don't, if like, you're not, if you can't show it like rigorously, like what kind of problems neural networks are good in solving? I think what I'm trying to look at is a symbolic thing, belong to the broader problem class of building mathematical foundations for reason. Reasoning itself have a lot of logic and discrete structures involved, system one versus system two. And how neural network forget to learn those things. I mean, even with the channel solved involved, it's basically a miracle. So I have been working with some colleagues. All those COT based things could be represent in us, you know, a superposition of logic. And maybe our theory can get somewhere from there. And that's something we are still on the work. But I end. We also want to work on practical side of reasoning. And I'm not saying that we have to necessarily wait for the whole theory to complete. I would say they just fulfill the both sides of my curiosity. And do you think actually do you want? No, no, no, no, keep going early. Okay. Okay, so I would try to push you in front of you. Do you think reasoning is actually like useful or something that you each some kind of like I'm told local minima that we that we have now and we can we can bypass it with more advanced models like why we actually need reasoning? Well, let's first of all define reason you which you call as reason is yeah. The O T. Reimpulsing already. Oh, what a specific kind of reasoning. So what do you mean by reasoning? So yeah, I thought it more like yeah, let's say our reasoning, what people call reasoning models, more chain of thoughts like all these groups of solutions. Do you think it's actually helpful? Not helpful, but like for the long term, do you think we we these are essential for better models or we can make better models without it? Well, I'm not sure about it. I don't know, Elling algorithm is truly truly that necessary. If you have good enough data, I think what we all know that frontier labs are very heavy on data. I think all the algorithm can be viewed as a sound sort of search in high dimensional data space. If you have designed a good architecture or good optimization algorithm, you start from a good inductive bias of search. So that's that warm up initialization that you really serve you well. But if you have really good sampling of data from your target distribution and keep running random sample even maybe that will also perhaps also bring you where you want to go. I mean, I don't have any theory on that. I just you know, because in my current industry role, I work for XTX market. My work involves a lot of large model training as well. I feel this is also a process of moving my belief away from highly crafted or specialized algorithm model design to more fundamental perblonsetting or data distribution or any more inclined towards screwing 19 the later that you really give us more performance boosts. That's probably what I can say. I wanted to join in on this because I remember a paper. I think it was from more than a year ago, it might have been like an ACL paper from late 2023 on this overarching question of getting symbolic representations from, you know, techniques with gradient descent. It's a ridiculously simple technique too. They were talking about oh, this is my paper. Not your paper. No, no, no, I'm talking about a paper related to this. Oh, yeah, I saw. No, this paper, I don't remember the title. But the idea was for any neural net, you can get very accurate, fully explainable, extremely cheap spam classifiers, right? And all they do is they prompt the model to generate like their own anti spam classifier. And so it will like generate keywords or some other pretty good heuristic that's super explainable and ultimately creates a set of if else's that make it basically a decision tree. And you can use that for, you know, figuring out if something spammer or not in the real world and they show an experiments that the models that are created are pretty good, right? So I'm intrigued by this idea of just asking models for symbolic representations. And I'm wondering, I obviously this is nowhere near as elegant as I think most of the stuff you're directly talking about, right? But I'm just curious what your thought is to this approach. Well, I think this is an interesting approach. Although I think if you're believing this approach is basically you're believing to how well the model alignment is doing, right? That's specific dimension. And if language model will be talking regardless, what do you ask? So if, well, so the data, the data quality you get from language model really depend on the quality of the question you make for them. But I think that's why I'm mailing L.M. papers as such harder to reproduce. But okay, going back, I think you're all told. Talking language model itself is definitely a very interesting research direction. And in a broader sense, I also call it a new world symbolic AI because it's generating symbolic language. And the language itself is subject to symbolic checkers, depending on what domain you are. So sure, this is a very interesting interface that my research group is also personally looking at. We have done works to show you to, don't work like asking language model to propose planning solutions and then send the planning solutions towards the security checkers if that's a domain subject to compliance. And use that feedback to run DPO type of thing to directly optimize the language model to making it more compliant. Then if the data is good, the question you ask is good and the base model is good, you will be able to get this positive data fly we're going on. And yeah, you can, no guarantee to my best knowledge, but yeah, you can get that work. I have a question. So what do you think about how it's connected to all the JEPA styling learning? Like you have some internal state that you want basically to separate between the generation process from the learning process. And in some sense, like, it's not a script, but like you need some state, like right internal state that you borrow over time, you can think of it a bit as like symbolic model, right? Well, is it the necessary question for anyone coming to your podcast? Yeah, yeah, yeah. This is our, we need to change our name to JEPA. Okay. The JEPA podcast, but I think, yeah, I personally read the JEPA papers. I had to list the one to, I think I will, three or four Yen's talks on JEPA to make sure I understand all those concepts. It's a beautiful algorithm. I, first I like the MPC concept, although model based predictive control. And I think Yen's idea on combining this with a with a JEPA representation is beautiful. I also think it's deeply connected with other things that we have now, now better or earlier in deep networks, in deep, in deep network dynamic, such as a coupon operator, which I'm also very interested in. I, however, do not think Lansing is a, you know, all functional fix to all the prevalence we have in deep learning, because every learning algorithm has to assume something. And that's a sacrifice you make to, to the real data. Again, this is the second time I mentioned, I say, I, I say about this code, every algorithm is wrong, some are useful. I definitely belong to JEPA useful algorithm. Do you think that it's going to remain the core architecture behind world models? Well, I don't, like, like, personally claim to be any world model expert, so I don't know if I'm a great person to answer this question. I just feel like, I, I, I just feel like I have said, once the data have come to a point, all algorithm will convert. and probably a believer of the great convergence theory in model and architecture. I think there are all different avenues we take to achieve the final truth, which is a data dependent. Or maybe that is the joint entropy of data distribution. Who knows, right? Yeah, do you think everything is just a data? And then we think like we will converge with good enough data will converge to the same type of solution whatever like algorithm we will use. Well, yes and no. I think this is like the argument we have between network and neural network expressiveness and the practical youthfulness. If we neural network is universal or pretty, if a person made, we know that from the early 90s. But if we already know that by then, what we're doing in the past 30 years, we develop a new neural network architecture, resonate all fast-to-art things and that just receive a test of 10-mile word. That's because although neural network can in theory learn everything, it does not have to happen in your experiment. Your experiment is the further decide, your experiment quality is the further decided by how stable your optimization could go, how well your hyperparameter choice should I choose in, how good your architecture scales with your analysis, how does your architecture favors the DDP or FSDD or other things. And I think those lottery-line choices has been mattering on a unpropeotional portion in affecting the deep learning progress. So every algorithm are more or less useful. That is my view. But some algorithm are more advantage in today's ecosystem. And we should honor that. So when I'm doing research especially now, I'm doing part of the part of my industry researcher. I'm less religious about the algorithm is right or wrong. But I do think there are more suitable for now and less or less suitable for now algorithm. Keep what we have. I hope that is more or less answering our question. Yeah, I agree. I think I agree. I think it's a good point. I just feel that it's really, in real time, it really matters. All the small tweaks and all the small algorithm that what we are going to do and how we are going to do it. But after it, after you achieve it, then it looks that no, you can come up with it. You can do that many other ways. Yes, yes. And it's quite surprising, right? When you achieve your part. Once you know how to do it, once you know the first way to do it, you will know the second third many other ways to do that immediately. Yes, that's indeed what I have. Yeah, what do you think? So one more question that's related to this. What do you think about synthetic data? I just so I think it was Alex, the Amkis, I think he is in Berkeley now. Professor at Berkeley that said that they are doing a lot of synthetic data sets, generation and they are working on different directions. But one thing that they found out is that if now you need to create a data set. Let's say that you need to create a data set with 1,000 questions. Okay? And you are using some LLM to label these questions. I don't know, deep seek R1 or something like that. Apparently, if you have, so let's assume that you have one data set that with 1,000 questions and one answer for each question. Okay? And if you compare it to a data set with 500 questions and two answers for each question, then it looks that the 500 questions data set is better or it gives you better performance when you train on this data set. It contains more information in some sense. And think about it like, "Undo, it's quite weird," because you have the first data set, you have more questions. And this meets between what information we have in the questions versus what information we have in the answers and how they interact with each other. It's something that I don't know a lot of work about it. So, yeah, what do you think in general about a synthetic data set and how you actually use it in real life? Okay, yeah, I can have two parts of answers to this. Regarding the, regarding the person between you have one have more questions or you want to have less questions for more answers. It's actually the first time I heard the example you say that I would look into more. But I think earlier, a few months earlier, my friend Professor Alex Dimarches of Berkeley, he was rising, I think he was posting about that paper, he's paper, I'm not sure, but he definitely has posted a paper saying that you can just train reinforce learning using one question. Just using one question, getting to the answer and then pushing the model again, could you do it a bit more differently? Could you do it more creatively? Could you do it with another way? And keep doing that and see the model, a lot of self-improvement on some small-based model as well. And there may be another paper related to this idea. I think I have C2. And that's quite interesting. And when I look at reinforce learning papers, I would like to connect back to our own life experience. I remember when I was a high schooler, I was practicing my master. I first look at, you know, I do a lot of master exam questions and try to make myself just practice myself. And as a my teacher, look at me and saying that, no, I don't want you to do this because that's a great way to overcome. I want you to work on this, select the book, and then keep doing this question once and do it again and see if you can solve it better and eventually better than the standard answer provided to you for reference. And that becomes what I followed. That didn't make me a good mathematician, but I think I have to see both sides of this practice. Sometimes a deeper tool is better than the diversity of the permanent look out to. I can see why this makes sense. It's just trying to, I think it's at least, I try to understand it as squeezing the salt value from the salt process. If you just look at the question, get an answer and give it up. It's like you are being served as a dish and you just give one test of it. You didn't really finish it. There are more intellectual value you can get from this question and that's how you can benefit from repeatedly chewing this question, testing this dish. So while I haven't looked deep into the paper, I think I can buy this idea. And I think this is an inspiring one. Second to the sensitive data question. Yeah, I can comment from my practice. I think synthetic data is actually a luxury. You only have in certain domains, right? Probably language and vision and maybe speech but I don't personally work on speech. Those domains have a commodity. We have all worked in these domains for years to make first discriminative models work good in the domains. Everything starts almost from a recognition type of thing, classification recognition. And we do other types of multiple discriminative questions, maybe labeling, maybe detection or segmentation, or so on. We have, and we made a sufficient progress on discriminative models and the bigger the larger datasets for them. For example, imagine that was a buger for not the buger for generation in the beginning, right? And even Lyon was a buger for captioning, not for today's generation in the very first stage actually. And only when the discriminative models has worked well enough, we went to the generating model stage. And we reused the high quality dataset we build in the discriminative stage. And it is also so natural from the machine learning commonsense perspective because the joint distribution is always the hardest to learn, right? It's the God of all distribution. And so that's a natural process. Only when you have the good generating model and learns the good joint distribution, we can start talking about luxury of synthetic data. So this is such a high bar that only a few lucky domains, including language and vision have. And I personally now working on one of the very unlucky domains called high frequency trading. Well, if anyone knows how to build generating model to generate the everyday stock market, I'm looking at please, if you are listening to the podcast now, I would like to talk to you. So you can make your own hedge found. And I'm sure that they can all say how important are things. And mind you, I'm asking about like the industry in general, right? Like how important is cabling to high frequency trading? Like for example, being close to those data exchange hubs where data like comes in on those undersea cables, I've heard that things like that and like A6 and FPGA's are very important because for high frequency trading, my understanding is that you have like in some cases nanoseconds of time to make decisions based on information. Uh, yeah, speed is very, speed is very important. But there's also no denying that generative AI revolutionizing this industry. That's a partially why you are seeing such an enthusiasm in both the exhibition hall or the workshop rooms in New York's last week. But Ravi, you are asking what I can say about this industry in general or user. Yeah, like what what types of problem using our interest in this industry in general. We don't have a lot of problem types, although each problem is very deeply, have to be very deeply studied. The most popular, I mean the most useful, the most common and useful, and perhaps the closest entry level quantum research question is time series forecasting, multi-round time series forecasting. You take input features from the market exchange, in high frequency case, you take input feature from market exchange that concludes the price, volume, transaction record, blah blah blah, as a metadata for each stock. You may want to look at the mailing stock together to build a so-called foundation model. But the goal is for each of the stock, you want to predict their price, or at least the direction, after a certain period. And that certain period called the price, you really decided by what execution strategy you use. So it's a very classically defined question. And the problem is the classical, why you call FX, why the price, X is the current price, and the past price, and other associated information. The research challenge here though is the data is extremely noisy. And as I said, I don't know how to best utilize the blessing of synthetic data, because we feel if you're working the high frequency field, you are either, we are not short of data. The exchange are sending us data every tick. The issue is we are short of high quality data. And there is no obvious way for me to tell, I mean, for me to tell whether today's stock market is better than yesterday. And note that here by quality, I'm not saying the data itself has mechanical error or transmission error. It's not, and we have great data team to help us ensure the data quality are as authentic as possible. The issue is the market itself is always dominated by noise. So that makes this pan-serious task specifically a bit like predicting noise from noise. And in fact, if you're working this field, you know the prediction reality is very much like zero correlation. And that's why typical quant is actually challenging for typical retailer investors to beat the market itself. For example, S&P is a very good and numerous neutral baseline and it's very hard to beat. And people all know that, right? We did so many efforts just to make our prediction accuracy to be a little above zero. And that a little margin is very small for all high frequency quant, but because this industry, people trade a lot every day and every year. The total trade we execute is an astronomy economy number. And with a simple, with a small margin, every trade, but to do that a lot of times, then the law of large numbers will be off-rand and materialize this probability into a number that are enough to cut into your checks. That's what we do. But do you think it's such a hard problem because at the end you don't have one solution, right? It's not like, it's not a deterministic, right? Like you don't have, like, I don't know, you don't know at the end, like it depends on the opinion or so of so many people, right? So it's really like even if you have all the information, let's assume that we have all the information, even in this case, it doesn't clear that you actually can, like, guess a good prediction, right? Of course, of course. Yeah. Yes, by definition, market is a multi-party game problem. And the number of participants is so large that it's impossible to have any analytical solution. In the specific case of high frequency trading, I would say this is a, I mean, all this has been less considered as a concern because the time you have to make your trading decision is such a short horizon that, well, let's put it this way. This is too short for anyone to float against you, if I can say this way. It's not necessarily true in recent years because there are more advanced things that people are doing anyway, but I prefer not talking about that. So yes, the game, the multi-party game concern you have was not a major concern, but now it's getting into real as well for some reason. And without going to a specific details about models, I think, like, at least from someone that looks on it from outside, it looks like, you know, there is kind of like a lot of companies are actually trying to build and like to become less conservative, right? Like many years these companies were really conservative about the models and actually, like, the models themselves, like, were really simple in the sense that, like, they didn't follow after the recent advantages in machine learning, right? And it looks, again, from the, like, really looking from outside, it looks like that a lot of companies are trying to take some some metals and take some new frontier ideas from machine learning to make better prediction. Do you think this is something that, like, you see also in the field, you think this is something that will become more and more, you think something like you will see that will go back to more classical methods? Well, I think the benefit of a generator AI for trading is real and because there are at least a few companies, including XTX, has delivered the benefit. And I think people would be living what they see for the trend. It's also real. I think Alan and I both see the what's going on in Europe's conference. And the finance and the entire finance industry is of course very excited about AI tools because Alan's thing that can help make profit is good. Different size of finance industry may have a different take. For example, some might be interested in leveraging LMS to automate their workflow. Some may want to use a foundation model to read the internet or dig into social network and find what we call alternative R-Files to, like, for example, sentiment about the company. This is a famous example. And some like us and a few other competitors are trying to build a better foundation model ourselves, but not about language. It's about the time series, multivariant time series and the big collection of them received from each change. So instead of symbolic language foundation model, like everybody else talking about, we're talking about a numerical continuous stream time series foundation model on time series and a few other things. So people have different target and different usage when they talk about AI. But one thing that I think we can make no mistake on is everybody yes, is trying to use AI. And I can also see AI is already delivering in many other firms. It's probably no stranger news, no stranger news anymore that there are banks and there are trading firms who are posting who already either already build the AI center or aggressively post the job post about where high remote AI researchers. So I'm also telling my students that if they're interested in finance, this is actually a very good year to jump into finance job market or to at least the fill how it's like. Yes, some of my students are interviewing as well. Regarding whether all of the more deliver that AI cannot predict, but I guess even we're looking to the tech market, we even cannot be sure which of those whether all of the LM company will deliver. Yeah. Finance are slower in catching the trend in general. The dynamics of finance is always lagged and the EMS moves the trend compared to what's going on in the tech field. I think some people will succeed. We already see successful examples. That's the most important thing. You can not convince finance people with papers. You have to convince them with annual returns. And that already happened. People were moving to the direction and I expected to see more companies to keep investing and sound them well succeed. And do you think so? So classically, let's say the average profile of the 22 finance was a graduate PhD math student. Do you think we will see more and more ML people, people that actually build models, AI models that know how to deal with models or we'll see some mix. What do you think will be the future and direction in the profiles? Yeah, that's a great question. Because I'm always, because I'm also hiring, I can show one of the models that we hire quantum researchers here. And that was not my credit, it's a team credit. We want to mathematicians who can write code. All engineers who know how to speak math. Maybe they're being the same thing. But I think I'm telling you that I think in general, the hiring preference of frontier financial AI confidence, if I would call in this terminology, I become more and more aligned with the frontier AI labs that people are talking about. It's pretty common for me to see people in our candidate pool to have compete offers from an open AI gym night. It actually happens from day to day. If I want to name one relatively unique preference to finance, it's perhaps we still hold our preference towards mathematicians and statistical foundations. Like, okay, you are still not get away from the problem like tossing one's old and coin, but what's the first time you see, I probably still see these questions somewhere. Not necessarily with us. But we want people who Dior can't out-sade because our data is very noisy, and you still need to do your research work. To even you go with a foundation model recipe, there are still a lot of research you have to do to bridge between the foundation model of period for clean structured language to the noisy less structured time series. And there are many mathematical transformations linking this bridge. So this is just one example. So we still want to people who have those condensers. Among my colleagues, there are top mathematicians, the top of physics, researchers, quantum researchers, and of course a computer scientist and a lecture engineer. Eventually nobody come out with a peer-to-degree in trading, right? There are just the mathematicians who can write a code that are trying to buy different disciplines. - Yeah, I agree. I'm not sure right there is all this debate recently. Like what will be now when you have all the AI tools that know how to write code, there is this debate, what is the best and the degree to actually learn and what you actually need to learn in these courses. Do we need to learn to know how to transform a matrix or not, any algebra and all these things. So personally, I believe that yes, I believe that like in the last 30 years, the best choice was to actually learn something like as as Matthew as you can, right? Like to when linear algebra, when like calculus, whatever you can, as much as you can, it would be better in the in the drug market, right? But it's not clear if this will be the case in the future, even for AI research, I could say pure AI ML research. - Well, that's why you got to get your plumbing license as soon as possible. - Well, that also depends on how soon you believe in the embedded agent, the embedded intelligence is gonna succeed, right? - Not even 100% safe. - Well, I still think physical, physically demanding blue collar work like that will survive a bit longer than other types of physically demanding work. Like for example, it's apparently easier to create loitering like munitions and replace soldiers than it is to do plumbing work. 'Cause we have a lot of the munitions stuff already. - Yeah. - And maybe I agree. And maybe we can also look further into areas where the either data availability is less sufficient. Our areas will prepare to be data war has been holds there and there is no much way to penetrate that. As in finance also belong to the last area. They have, this is very hard for a generic public model to penetrate the private knowledge war. And the private knowledge accumulated and finance industry in the last decades or decades is definitely valuable. And you cannot easily override that with public knowledge. That's my personal opinion. - So what do you think about all these competitions that let the item, Gemini, Chagypety, Deep Sick, they give them a free trading account and let them trade over two months? - That's an interesting experiment. I will not put my personal money in it. - Okay. Okay, no, that's a, I agree, I agree. Anyway, do we have anything to have anything there? Elsa, do you want to talk about, or do we comment about? - Well, not really. I think we'll have touched a million interesting points. Actually, more intense than I thought. Still, better than the review, too, I get, which I appreciate. The last word I want to comment is that maybe I can add a last work comment and to just the better fit the industry had, I'm worried. I personally think, I mean, fun, when I did my PhD 10 years ago, I didn't actually give much of consideration to the finance industry because I was on there the same stereotype of perhaps many people have. This is a linear regression, last one, the different. Paid very well, but I want to do by that time, computer vision, all the machine learning stuff. But now I can see this field is embracing a real change because the pressure of real market success, but a heap, and those two are very different, right? You can enter a field where you hit a lot of heap, but then you find it's very difficult to find one company that actually makes money and rolling the flywheel up. But finance already proved a concept. That means it already passes in fanhood and it enters a maturity time of using AI. I firmly believe this, we will see the explosion of AI in finance because of someone proved it will make money. The research of Perblon here is amazingly rich and amazingly difficult. And I'm saying that in a positive thought, I was personally facing choices between going to for example, LAM group and the current firm I'm working for. And I was debated a bit and then I made my current choice because I feel if I join LAM Foundation Lab, I think I'm a reasonable smart guy. I can learn how to do pre-training or post-training if someone guides me. But joining my current firm will allow you to do things, allow me to feel a few thousands of people to know how to do right in this world. And because of the unique data access, the M-POT GPU power and the talent intensity you work with. And I think I like being unique, I always like being unique in my research journey. And I think in that sense, I'm very happy with my current firm. So the entire industry will open more opportunities and I believe some of them will succeed. For AI researchers, actually, new grad, when you look at job opportunities, I would offer the advice, please, look at this industry more seriously because we need people to do research. I think we pay people pretty reasonable, you're unreasonable, well, maybe compared to today's tech industry, we have a better work life balance. That's all the word I want to say. Yeah, great summer. I think we should talk about actually work balance, work life balance, once a day, because the卫ter is a huge, some companies really emphasize and really know how to do these work life balance. And some of them it's terrible. And people think that if you will work much harder and if people will work 20 hours a day, they will get better results. Yeah, I have a small sample size of my graduate PhD student. Some are living a happy life. Some basically disappear from social life. So I can open up. Yes, yes. And in the bottom line, it doesn't, mostly the time it doesn't matter. Like for the final product or later, the final result. So, yeah, but this is, but it's to do it. Yeah, but it's today's culture and it's a great for young guys and students to be motivated. So I wish them all the best. I wish them the very thing schedule all work well for them. Yeah, and it's out there, and specifies specific names, right? Anyway, that's all. Thank you so much. Thank you. Yeah, so put your stuff on this. All right, thank you. This is a great conversation. I actually feel more energized after talking to you guys. And I will probably surreal the paper anyway because I really like that paper. Yeah, yeah. Thank you so much, everyone. And see you next time. See you and thanks everyone. Bye. Bye. [MUSIC PLAYING] [MUSIC PLAYING] [MUSIC PLAYING] [MUSIC PLAYING] [MUSIC PLAYING] [BLANK_AUDIO]

Podcast Summary

Key Points:

  1. The podcast discusses NeurIPS 2024, with mixed feelings about its size, quality, and mix of researchers, VCs, and industry presence.
  2. Atlas highlights the value of workshops over main conference sessions for more pure science and frontier ideas.
  3. Alan notes frustration with the conference app (lack of HOOVA) and the lower ratio of actual researchers to attendees.
  4. Both guests appreciate the networking opportunities, but critique the increasing commercialization and quality issues in AI conferences.
  5. Atlas’s research focuses on low-dimensionality in neural networks, including sparsity, low rank, and symbolic compression.
  6. A key recent work is a theory paper on how neural networks can learn symbolic equations via gradient descent.
  7. The discussion touches on unsupervised learning, priors in human cognition, and the goal of compressing neural networks into symbolic, human-readable knowledge.

Summary:

In this episode of the Information Botanical Podcast, hosts Robert and Alan welcome Atlas, a faculty member at UT Austin and research director at XTX, to discuss NeurIPS 2024 and broader AI research themes. Atlas describes the conference as a dual role—presenting academic papers and staffing an XTX booth—which made it busy but rewarding, especially for interactions at the intersection of AI and finance. Alan shares mixed feelings: he enjoyed the LLM sampling research and San Diego’s atmosphere, but was frustrated by the poorly received conference app and the high proportion of non-researchers, including VCs, which diluted scientific discourse.

Atlas defends VCs as technically savvy, noting one who had read his papers, and emphasizes the value of workshops for pure science, citing examples like the pluralism workshop with Ted Chang. The conversation shifts to Atlas’s research on low-dimensionality, a theme spanning his work from compressive sensing to pruning, low-rank methods, and symbolic learning. He highlights a recent theory paper on how neural networks can learn symbolic equations via gradient descent, arguing that the ultimate compression is converting neural knowledge into symbolic, human-readable form.

Alan connects this to unsupervised learning and human cognition, though Atlas cautions against over-extrapolating biological priors. The discussion underscores ongoing tensions between commercialization and scientific integrity in AI conferences, while celebrating the enduring value of low-dimensional representations and symbolic reasoning.

FAQs

Atlas focuses on low-dimensionality, including sparsity, low rank, and manifolds, and has recently explored symbolic neural networks to compress knowledge into human-readable forms.

His favorite recent paper is a theory paper on how neural networks can learn symbolic equations via gradient descent, aiming to bridge continuous dynamics and discrete structural learning.

Atlas prefers workshops for their purer science and more open sharing of half-baked ideas, as they are less influenced by incentives like recruitment or funding.

Alan was frustrated that the conference didn't use HOOVA, the web app, and found the replacement poorly received and lacking features, which made him feel disconnected.

Atlas respects VCs for being technically prepared, often reading papers daily, and helping to select good investments, which he sees as a challenging task similar to paper reviewing.

He considers compressing a neural network into a non-neural network, such as symbolic knowledge or human-readable language, as the ultimate compression.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.