This podcast episode traces the origins and evolution of optimal control theory, beginning with its philosophical core: determining the best actions over time. It identifies the 1696 Brachistochrone problem—finding the curve of fastest descent under gravity—posed by Johann Bernoulli as a pivotal moment. This challenge moved beyond static geometry to optimize a dynamic process (time) under physical laws, effectively birthing the field. The story unfolds through historical narratives involving key mathematicians like the Bernoulli family, Newton, Leibniz, Euler, and Lagrange, highlighting their intellectual rivalries, such as the bitter calculus priority dispute between Newton and Leibniz. The episode explains how these early efforts in the calculus of variations systematically developed methods for optimizing functionals, setting the stage for 20th-century revolutions like Pontryagin's maximum principle and dynamic programming. It frames the entire history as a sustained quest to formalize necessary conditions for optimality, connecting a 17th-century puzzle directly to the mathematical foundations of modern engineering, physics, and computation.
[Music] Hello and welcome to InControl, the first podcast on control theory. Here we discuss the science of feedback, precision-making, artificial intelligence and much more. [Music] I'm your host Alberto Paduan, live from a recording studio in Vancouver. My voice today decided to take a day off, so please bear with me today. As always, quick thanks to our sponsor, the National Center of Competence in Research on Dependable Ubiquitous Automation. So every day we solve tiny control problems without even noticing. How do you tilt your coffee just enough to avoid a spill, or perhaps a bit more consequentially, how can you shave minutes of the daily commute? What is the fastest route to work? Now raise the stakes. Imagine landing on a moon with your spacecraft. How should the trusters fire to minimize fuel while ensuring a precise touchdown? Beneath all these questions from the mundane to the spectacular lies the same timeless question. What is the best way to act over time? This deceptively simple question gave birth to what we now call "Today Optimal Control Theory", a discipline whose roots stretch back for more than 300 years. Long before the term "optimal control existed", mathematicians were already grappling with problems of best paths and extreme values. Their ideas laid the foundations for a field that would eventually provide the mathematical foundations of many important developments in physics, engineering and computation. At its core, optimal control is the art of making decisions over time to achieve the best possible outcome. So in this episode we'll take a journey through centuries of discovery. From a 17th century puzzle about a falling bid, through the 18th and 19th century breakthroughs in the calculus of variations, to the 20th century revolutions sparse by contraggins, maximum principle, Belman's dynamic programming, and Kalman's state space control theory. Along the way we'll meet a cast of brilliant and sometimes a bit eccentric characters, like the Bernoulli brothers, Newton, Euler, Lagrange, Hamilton, Pontriaggin, Belman, Kalman and many more. And we will explore how their insights, rivalries and the worlds they lived in shaped the development of optimal control. I should say from the outset that this episode draws inspiration from a remarkable paper by Hector Susman and Jan Villemmes, titled "300 Years of Optimal Control". From the Barquistocrone to the maximum principle. Their work serves as both a scholarly roadmap and a reminder of how far this field has come, and how far it might still go. Our story begins with a public challenge issued in the Acta eroditorum in June 1696 by the Swiss mathematician Johann Bernoulli, where he derled and I quote "the most brilliant mathematicians in the world to solve a new problem". His problem, dissettably simple at the surface, was this. Find the curve of fastest descent between two points under gravity. Not the shortest distance, but the shortest in time. This would become known as the "brachistocrone problem" from the Greek "brachis" meaning "short" and "chronos" meaning "time". It's actually interesting to quote the whole piece directly from Acta eroditorum, which is available online and of course I'll provide a link in the description. The piece says, and it's delightfully simple, it's delightfully short and says the following. Invitation to all mathematicians to solve a new problem. If in a vertical plane two points A and B are given, then it is required to specify the orbit A and B of the movable point M, along which it's starting from A and under the influence of its own weight, arrives a B in the shortest possible time. And just a bit later, Johann Bernoulli says the following. In order to avoid a hasty conclusion, it should be remarked that the straight line is certainly the line of shortest distance between A and B. But it is not the one which is traveling the shortest time. However, the curve A and B, which I shall divulge if by the end of this year nobody else has found it, is very well known among geometers. At Leibniz's suggestion Bernoulli extended the deadline to Easter 1697, on January 1st of that year, he issued a new global announcement boldly addressed again to an "the sharpest mathematical minds of the globe". Johann Bernoulli wasn't just a brilliant mathematician, he was part of a dynasty. The Bernoulli's were a Protestant family originally from Antwerp in what is now Belgium. In 1583, fleeing the religious suppression of the Spanish Aspergs, they left Antwerp and after spending some time in Frankfurt eventually settled in Basel, Switzerland, where their intellectual lineage would flourish. Across three generations, the Bernoulli family produced no fewer than eight notable mathematicians, enough to fill a faculty roster of their own. The most prominent of the clan was certainly Jacob, Johann's older brother. Then of course there was Johann himself and Johann Son Daniel, who would become a giant in physics and aerodynamics. Jacob was a pioneer of probability theory, most of you probably know the Bernoulli distribution, which bears his name, and Daniel later discovered the principle in fluid mechanics that still carries his name, Bernoulli's law. But when Johann came to age, the world of mathematics was undergoing a transformation. In 1684 Leibniz published his first article on differential calculus in the Aqta eroditorum, and the Bernoulli brothers were among the first to master his methods. In 1691, Johann used the new calculus to determine the shape of a hanging chain, the so-called "catenary", a success that won him early acclaim. That same year, a young and ambitious Johann was hired by the Marquide Loupital, a French nobleman and one of the foremost mathematicians of the era. Johann was tuturium in the mysteries of the new calculus, but there was a catch. Loupital paid well, but their contract gave him the right to claim credit for any discoveries Johann made during their sessions. And claimed them he did. Years later, Johann would insist, bitterly, that it was the true originator of Loupital rule, something that we used every day, or those who study engineering and mathematics use every day for limits. And this was published in Annalise de Anténimo Petit. His contemporaries were skeptical of this attribution, perhaps due to Johann's combative reputation, but in 1922, lecture notes from their sessions were uncovered. And they appeared to support Johann's claim. To say that Johann Bernoulli was difficult as an understatement, he argued constantly with colleagues, complained about students and clashed with university authorities. In 1695, he wrote a letter to Leibniz, who suggested him to join the faculty at Gröningen, and as I mentioned shortly after arriving in Gröningen, to take a professorship, Johann fumed, and I quote, "I have not met any of the practitioners of algebra, which you consider present in Holland. To the country, I have not had the honor of meeting a single person who would even deserve to be called a mediocre mathematician. Even on to complain that teaching slowed him down, famously saying that, and again I quote, "The more progress the students make, the less progress I make." The rivalry between Johann and his brother, Jacob, was also notorious, but perhaps even more remarkable was Johann's treatment of his own son, Daniel. Johann accused him unjustly of plagiarism, and even threw him out of the house when Daniel won a French Academy of Sciences prize that Johann himself had hoped to win. Despite the venom, Daniel remained respectful, though he confined his frustrations to his friend, Leonor Euler, whom we will encounter later in this episode, and who was then a student of Johann's and later a close colleague of Daniels in St. Petersburg. Even symbolic honors didn't ease tensions, when Johann and Jacob were jointly elected to the Paris Academy in 1699, it wasn't the condition that they promised to stop arguing, a condition both promptly ignored. And yet, for all the conflict, Johann's legacy endures, at the University of Gruningen, where he once taught a stained glass window that picks Daniel Bernoulli clutching his father's robe while Leon proudly displays his solution of the Brachistokonom problem, the problem that we started with, and the problem that launched a new era of mathematical thinking at the beginning of what we now called optimal control. This is at least what actor Susman and Jan Williams are giving that paper. So let us go back to that challenge. It may be evident that at this point, with this challenge, Johann Bernoulli wasn't really seeking a solution, he already knew it. What he really wanted, though, was to provoke his peers, or possibly expose their weaknesses, or perhaps simply assert his own genius. He was a characteristically aggressive move from a man who was both a brilliant mathematician, and also a very difficult person. Jan was renowned not only as I say for his mathematical prowess, but also for his legendary arrogance and volcanic temper. And despite quarreling with almost anyone, even with his own family, among those he resented the most, was certainly Sir Isaac Newton. Now, there are few intellectual minds in the human history that can rival Newton's. Some credit to his amazing genius, the foundations of mechanics, establishing the groundbreaking theories in inertia and gravity, optical discoveries such as observing color is a property of light, studies on the speed of sound, and groundbreaking works in natural science and philosophy, as well as huge leaps forward in the field of mathematics. However, one cannot please everyone, and one of Newton's most aggressive critics was precisely Johann Bernoulli, who in 1696 went particularly far in trying to dull the glowing admiration of Newton. Jan Bernoulli was not a fan of Newton due to a debate over the origin of calculus, between him and the mathematician Gottfried Leignitz, whom we mentioned earlier before. Newton claimed that he had been working on a form of calculus in 1666, but did not actually publish it until decades after the fact in 1693 and again in 1704. Leignitz developed his own form of calculus in 1674, and a 1696 publication about Leignitz's method mentioned Newton's work the Principia, as being nearly all about this calculus. Newton, for his part, was hardly a neutral observing the debate by the Herliden 1700s, yet become president of the Royal Society, a position that gave him immense influence over the scientific establishment in England, and of course, used his power. In 1712, the society released a formal report investigating the question of who had truly invented calculus. The report concluded unsurprisingly that Newton had priority and that Leignitz had likely plagiarized his ideas. What many didn't know at the time, though, was that Newton had effectively written the report himself, anonymously. The Royal Society's committee was little more than a formality. Newton curated the evidence, framed the narrative, and even edited the final text. It was one of the most infamous acts of scientific self-promotion in history, and it cast a long shadow over Leignitz's reputation for years. But Newton didn't stop there. He also used his control over the philosophical transactions, the Royal Society's flagship journal, to shape what did and didn't get published. He also delayed and discouraged and outright blocked some submissions by continental mathematicians aligned with Leignitz. In a particularly bitter episode, Newton refused to publish a paper of Leignitz's in the transactions, claiming it was either redundant or unworthy. The rejection sent shockwaves across Europe and deepened the sense that English and continental mathematics were drifting into open war. Extensions escalated between Newton and Leignitz over the true origins of calculus, Johann Bernoulli emerged as one of Leignitz's fiercest allies. He used Leignitz's own journal, the Aqta eroditorum, to pose the Brachistokron challenge to the intellectual world, believing that it could be only solved by someone who, and I quote, "truly understood calculus." Over the next months, no fewer than six mathematicians responded, and not just any six. The respondents read like a who's who of the eras mathematicians. Alongside Johann's own solution came one from Leignitz, who called the problem, and I quote, "splended" and worked out his answer in a letter dated June 16 of that same year. His older brother, Jacob Bernoulli, submitted a solution as well, despite their bitter feud, as did the philosopher mathematician Aaron Fried, Valter von Schirnaus, and the Marquid de Loupital, best known today for the rule that bears his name, as we mentioned before. But it was the final contributor who gave this story a mythical punchline. In February 1697, the royal society received a solution, anonymous, elegantly ters, and presented without proof. It was published in the philosophical transactions soon after. Yet, to Johann Bernoulli, the author's identity was unmistakable. He read it and reportedly uttered just five words in Latin, ex-Ungwe leonin, which translates as, "I can recognize the lion by his claws." Of course, it was Isaac Newton. Johann's own solution appeared shortly after, in May 1697, issued by the acta eroditorum, the journal where he had first published the challenge. By the way, all issues of the acta eroditorum that we mentioned so far are available online, and I will provide a link in the description. The issue also could include that Jacob's answer, reprinted Newton's anonymous submission, and featured the responses from Chimnaus and Opetal. There was even a note from Leibniz himself. It chose not to republish his own solution, remarking that it was so similar to Johann's, it would be redundant. He did, however, offer speculative lists of those he believed capable of solving the problem. L'Opetal, Huygens, had he been still alive, Hadne, had he not abandoned mathematics, and Newton, as if Leibniz rallied it, it would take the trouble. It is worth mentioning that Bernoulli initially believed incorrectly that the "Pracistocrine problem" was a new discovery. Leibniz, however, was aware that this was not the case. As early as 1638, Galileo Galilei had discussed the problem in his book Two New Sciences, and had even proposed a solution, which he thought was a circular arc. Galileo had, in fact, shown correctly that a circular arc always, outperforms a straight line, except in a trivial case. Bernoulli took Galileo's errors, his belief that the catenary was a parabola in that the "Pracistocrine" was a circle, as strong evidence of the superiority of differential calculus, or "nover methodus" as it was then called. He was delighted to discover that the true solution of the "Pracistocrine" problem is, in fact, the cycloid. This is another curve that already been introduced and named by Galileo himself, because of its connection to the circle. Oigons later identified another extraordinary property of the cycloid. It is the only curve along which a body moving under gravity oscillates with a period independent of its starting point. Contrary to Galileo's belief, a circular path possesses this property only approximately, since the oscillation period of a pendulum depends on its amplitude. For this reason, Oigons called the cycloid also the "tautocrine" from the Greek "equal" and "time" again. Bernoulli was both astonished and somewhat perplexed by the fact that the cycloid serves simultaneously as the "Pracistocrine" and the "tautocrine". Two seemingly distinct properties concerning the time taken by a fallen body ultimately led to the same curve. From this, he concluded, and I quote, that "nature tends to organize phenomena in the simplest possible way". As in this case, by assigning two properties to a single curve. Bernoulli's "Bracistocrine" challenge of 1696 is cited in the paper by actor Susman and Jan Dilems as the "birth of optimal control". Insolving mathematicians that to invent new methods to optimize a functional, travel time, essentially creating the calculus of variations. Setting the stage for future mathematical developments, championed by luminaries like Euler and Lagrange. Now, why should the "Bracistocrine" problem mark the "birth of optimal control" and not merely be considered as yet another milestone in the calculus of variations? After all, similar problems had been studied since antiquity. For example, the Greeks had already solved what may be the oldest of them all, determining the shortest path between two points, a straight line. They also grappled with more subtle geometric challenges, such as the infamous "isoperimetric problem", also known as "Dido's problem", which asks for the playing curve of a given length that encloses the largest possible area. The Greeks knew the answer was a circle, though it wasn't rigorously proved until the 19th century. Others like "Hero of Alexandria" used geometric reasoning to explain optical phenomena. In his catopteryx, "Hero" argued that when a light ray reflects off a mirror, it follows a path of shortest distance, and thus shortest time. Fermat, in the 1600s, later elevated this idea to a general principle. Light rays follow the past of least time. His insight would go on to explain not just reflection, but also refraction, and as we'll see, it deeply influenced also Bernoulli's thinking on the "Bracistocrine". Yet, none of these earlier problems explicitly framed time-minimizing paths in terms of dynamic behavior, of moving bodies governed by physical laws. That's where Bernoulli's challenge was different. In his 1696 problem, it was not just about geometry or length, but about time, motion, and control. Susman and Williams argue that Bernoulli's problem, as posed in the actire de Toroom, is a true minimum time problem, of the kind that, in my quote, is studied today in optimal control theory. What needed to be minimized wasn't distance, energy, or curvature, but time itself. What makes the problem deeply modern is that all the interesting structure comes not from the cost, but from the dynamics, the way motion unfolds under the force of gravity. The challenge lies not in choosing a shape or a geometric figure in isolation, but in selecting a path that respects physical laws while achieving a temporal optimum. It's not a static curve problem. It's a question about evolution in time under constraints. And that, fundamentally, is what defines optimal control theory. As the authors emphasize, in a quote again, "a large part of the subsequent history of the calculus of variations can be best understood as the search for the simplest and most general statement of the necessary conditions for optimality, a search that would eventually culminate in the maximum principle, the central tool of modern optimal control theory today." But more on that later. So after the Brachisto-Cron challenge, the search for optimal curves became a central theme in mathematics. What began as a clever puzzle, posed by Johann Bernoulli, quickly evolved into an entire branch of mathematical thought. And so in the 18th century, two towering figures, Leonardo Euler and Josef Lleith Lagrangeur, Stuchon had been elegant, but isolated problem solving, and transformed it into a systematic, powerful framework, the calculus of variations. Euler is one of the most prolific scientists in history. He was a Swiss prodigy born in Basel, studied at the University of Basel under Johann Bernoulli, as we mentioned before, who recognized this extraordinary talent early on and gave him weekly private lessons. By the 1730s, Euler was already tackling classical variation on problems like the isopereometric challenge. But his landmark contribution came in 1744, when he famously published the "methodus in Venni and Ilina Skurvas, Maximimimimi, the "properate godendess", or the method for finding curve, lines that possess maximum or minimum properties. This monumental text laid out a general method for finding extreme of functionals, mathematical objects that assign a number to a function, typically by integrating it over some domain, and introduce what we now call today the Euler Lagrange equations. Namely, a differential equation stating that the time derivative of the partial derivative of a special function, which we now call today the Lagrangian, with respect to velocity, must equal its partial derivative with respect to position. In other words, instead of optimizing simple quantities like numbers or points in space, Euler showed how to optimize entire functions, curves, paths, or trajectories of a system, that minimize or maximize some overall quantity. Like time it takes to travel, the total length of a path, or the energy of a mechanical system. The Euler Lagrange equation provided a necessary condition that any function must satisfy in order to be optimum. A kind of mathematical fingerprint that singled out the best path from all possible ones. Euler derived it by analyzing how a small change in the path affected the total value of the function, and identifying the condition under which such a change would leave the function. That is, when the path was truly optimum. He applied this framework to a wide-raiging problems in mathematics and physics. For example, determining the shapes of curves under tension, such as elastic rods or hanging chains, but also modeling geodesics, that is, locally shortest paths between two points, on curved surfaces, and calculating orbits and trajectories in mechanics. In each case, the physical situation could be reduced to minimizing or maximizing a certain integral, with the Euler Lagrange equation serving as a key tool for uncovering the solution. For the first time, there was a general analytic method for answering a question like, "What is the best path into a concrete mathematical equation?" One that could be solved, analyzed, and applied across physics, engineering, and geometry. It was a landmark shift. Euler had transformed optimization into a calculus, giving rise to an entire new language for describing the behavior of systems over time. And yet, a deeper revolution was still to come. While Euler gave variational problems their first general form, it was a young mathematician from Turin, who would revolutionize their method of solution. Giuseppe Ludovico de la Grange Tournienne, later known as Joseph Luix La Grange, was born in 1736 in the King of Sardinian. He was the eldest of 11 children and originally steered towards a legal career like his father. La Grange's father, Giuseppe Francesco Ludovico, was a doctorate in law at the University of Turin. While his mother was the only child of a rich doctor of Cambiano in the countryside of Turin. His father, who had the charge of the King's military chest and was treasurer of the Office of Public Works and Fertifications in Turin, should have maintained a good social position. And wealth, but before his son grew up, he had lost most of his property in speculations. A career as a lawyer was planned out for La Grange by his father, and certainly La Grange seems to have accepted this willingly. He studied at the University of Turin, and his favorite subject was classical Latin. At first, he had no great enthusiasm for mathematics, finding Greek geometry rather dull. It was not until he was 17 that he showed any taste for mathematics. His interest in the subject being first excited by a paper by Enbon Allé from 1693, which he came across by accident. Alone and unated, he threw himself into mathematical studies. At the end of that year, he was already an accomplished mathematician. Encouraged by the physicist Becaria, he began teaching himself mathematics and secret. Pains takingly transcribing the works of Wolf, Dalambeer, Bernoulli and Euler, whose mechanica he would later describe as a kind of an "university course at a distance". At 19, he was already a professor at the Royal Artillery School of Turin, and had helped found what would become today, the Academy of Sciences. He was in this context on August 12, 1755, that Lagrange wrote, "And now legendary letter, I will provide of course a link in the description to Euler." Inclused was a short appendix that would quite literally change the direction of Euler's thinking. In the letter Lagrange proposed a purely analytic method for solving variational problems, one that sidestep geometric reasoning entirely. He suggested perturbing candidate curves ever so slightly, analyzing how the total integral would change, and using that to identify optimality conditions. He introduced a new idea of the Lagrangian function, which in mechanics would often be the difference between the kinetic and potential energy of the system, and an analytic condition to generate Euler's equations almost mechanically. Euler immediately dropped his own path-leg geometric method in favor of Lagrange's analytic one, and even coined the term "calculus of variations" to describe this new field. I'm going to quote now from a piece that Euler himself wrote about Lagrange. In the summary of his first paper using the calculus of variations, Euler explicitly says, "Even though the author of this, Euler had meditated a long time and had revealed to friends his desire, yet the glory of first discovery was reserved to the very penetrating geometry of Turing, Lagrange, who, having used analysis alone, has clearly attained the same solution which the author had deduced by geometrical considerations. It is clear that the relationship between Euler and Lagrange was one of admiration, a mutual enrichment. In 1766, on the recommendation of Euler and Dalambeir, Lagrange was invited to Berlin to succeed Euler himself at the Prussian Academy of Sciences. He spent the next two decades there producing groundbreaking work in mechanics, analysis, and celestial dynamics. After the death of Frederick the Great, he accepted an invitation to Paris, where he had become a central figure in the scientific appeal of the French Revolution. By 1762, Lagrange had already published his results in full, laying the analytic foundations that would become standard later. His approach turned the once fragmented problem of finding optimal paths into a unified symbolic framework. And with these calculus of variations was no longer just a mathematical curiosity. It became a fully flage discipline capable of expressing dynamics, constraints, and optimization in one coherent language. Despite his quiet, the manner and bouts of depression, Lagrange rose to become one of the most celebrated intellectuals of his time. He helped develop the metric system, taught it the Ecole Polytechnique, and became an advisor to Napoleon himself, who granted him a title and a seat in the French Senate. Through it all, Lagrange remained committed to analytic purity. His magnum opus Mechanic Analytique, published in 1788, elegantly-reformulated mechanics as a branch of pure analysis. I worked that directly anticipated both Hamiltonian mechanics and modern control theory. It opens with the now-famous declaration that, and I quote, "One will not find any diagrams in these pages, only equations, a manifesto for the analytical power of the calculus of variations." Meanwhile, Euler, though fully blind in his final years, never stopped working. He dictated papers at Brax taking pace, returning to St. Petersburg, and producing some of his most influential work from memory alone. By the time of his death, he had authored over 800 publications. The work of Euler and Lagrange would soon further refine by Adrian Marie Legendre. A French mathematician known for his precision and rigor, whose work laid foundational stones across number theory, statistics, and mechanics. In 1786, Legendre introduced the second order condition now known as the "Legendre condition", a way to test whether a stationary curve truly corresponds to a minimum. In the calculus of variations, a variation itself, refers to a small, imagined change in the shape or form of a candidate curve. Like gently deforming a trajectory to see how the outcome changes. The first order variations help identify stationary curves, those where the change in the quantity being optimized, the function is zero. But not every stationary curve is a minimum, right? Just like we have infinite dimensional optimization. In principle, some could be maxima, and some could be saddle points. So, Legendre's insight was to look at the second order variations, how the outcome changes when you consider not just the immediate effect of a small change, but the "second order effects of change in the curve, that is capturing changes in its curvature in the space of all possible variations. If this second variation is positive for all admissible perturbations, it confirms that the candidate path truly yields a minimum, analogous to what happens in finite dimensional optimization. So, Legendre's criterion thus became another necessary condition for optimality, complementing the Euler Lagrange equation with a diagnostic tool to rule out false solutions. Legendre also introduced the now-familiar partial notation for partial derivatives, a symbolic innovation that would help formalize and streamline the calculus of variations in years to come. By the early 1800s, the stage was set for a radical rethinking of mechanics. Into this world, stepped William Ron Hamilton, the brilliant Irish prodigy who grew up almost as a Renaissance man, fluent in several languages before the age 10, winning arithmetic contests with visited geniuses. A detailed account actually of his life can be found in the paper "The Life and Early Work of Sir William Ron Hamilton" in the script of Mathematica, of course I will provide a link in the description. But Hamilton's life was not all equations, as a young man he fell hopelessly in love with Catherine Disney, an unrequited romance that after she married another plunged them into despair and poetry. He took refuge in literally inverse, seamlessly reminding visitors, including the poet Worseworth, that mathematics like poetry was a creative art form. Worseworth, ever the champion of imagination, gently scolded him, warning that "Science apply only to the material uses of life, wage war with extinguished imagination." These personal dramas, love, poetry, and friendship with a romantic poet, laughed Hamilton with a profound sense that nature's laws could be expressed in elegant, almost-litical form. And Hamilton found that form in what is now called Hamilton's Principle or the Principle of Stationary Action. In 1834, he announced a new variational approach to mechanics. The actual trajectory of a physical system between two states is the one that makes a certain quantity, the action, stationary, with respect to nearby paths. In other words, nature chooses between quotes the path for which the integral of kinetic-minus potential energy is extremal rather than reacting to forces instant by instant. The result was a profound shift from Newton's forces to a globe of energy-based view. In Hamilton's elegant formulation, the world could be described by a single function, the Hamiltonian, and the motion by paired first-order equations in generalized position and momentum. As mentioned, Hamilton's central idea was to reformulate mechanics through what became known as his Principle of Stationary Action. The notion that a physical system evolves along a path that makes a particular quantity, the action, stationary, typically minimized. At the heart of this formulation was his introduction of conjugate momentum variables, what we now call today also co-states, which allowed him to replace Newton's second-order differential equations with a pair of first-order equations. His insight was to package all the systems information into a single-scalar function, now called the Hamiltonian, whose derivatives give the time evolution of positions and momenta. In other words, it turned dynamics into geometry. In Hamilton's picture, every state variable has a corresponding conjugate momentum, and both evolve together according to the Hamiltonian's equations. As we shall see, Pontriagian would later discover a nearly identical structure for the state and co-state in optimal control, essentially re-imagine Hamilton's language in the service of optimization. As Susman and William Spurid, Hamilton was, in fact, a great contributor, probably the greatest single contributor of all time, through the calculus of variations. In today's language, his introduction of co-state-like variables foreshadows the adjoint variables that show up in more than optimal control. His elegance can be seen in the fact that when we set up a cos-minimization problem, the necessary conditions and that being Hamiltonian equations almost identical to what Hamilton wrote down in 1834. Hamilton went further and defined what he called the characteristic or principal function, essentially the minimum value of the action as a function of the initial and final states. This was a bold move. Imagine that instead of describing the optimal path itself, you describe how good that best path is. Hamilton noted that if you knew this function, you could recover all the details of the optimal paths by differentiation. In modern terms, this idea is exactly the ancestor of the value function in dynamic programming. Susman and William's describe how Hamilton's central idea was to regard the minimum value of the action as a function of the endpoints. And how difficult it was for his contemporaries to carry out his plan. But today's vision is recognized as the core of the dynamic programming. And again, quoting from Susman and William's, with the rise of optimal control, Hamilton's plan to develop all the properties of extremals from the characteristic function has become one of the best and most widely used tools. Under the name of dynamic programming. Hamilton was also a pioneer in optics, introduced what he called the characteristic function, essentially the action to describe light rays, drawing a remarkable parallel between the behavior of light and mechanics of particles. He showed that in conservative systems, energy remains constant and the action accumulated between two configurations, is encoded in the constant energy characteristic function. Incredibly, he predicted phenomena like conical refraction purely from theory and later confirmed them experimentally. But beyond those tangible achievements, his last in legacy was the radical idea that the true motion of any physical system is the one that extremizes action. In his characteristically terse and economical prose, Hamilton laid out results that would take decades to fully interpret revealing a rich trove of geometric and physical insight. His 1834 papers were difficult to penetrate, but they introduced a new language for dynamics. One that fused geometry, energy and optimization. Indeed, Susman and Williams even argued that if Hamilton had formulated his equations using the control as an explicit variable, he might have prefigured the optimal control Hamiltonian method a century earlier. Not content with Hamilton's breakthroughs, another German mathematician, Karl Gustav Jakob Jakobi, began exploring the same territory from a different angle. In the early 30s, Jakobi took Hamilton's idea and involved them into an even more general framework. He formulated what we now call today the Hamilton-Jakobi equation, a first-order partial differential equation for a function, the action, whose solutions encode all the dynamics of the system. In physics, this is a landmark moment. Hamilton-Jakobi theory provides yet another alternative to Newton's laws, one in which the motion of a particle could be seen as propagating along wave fronts of constant action. In fact, Hamilton-Jakobi equation would later become an important conceptual bridge to quantum mechanics through its resemblance to Schrodinger's equation. But Jakobi's contribution were also deeply mathematical. He showed that the Hamilton-Jakobi equation is equivalent to finding extremals in the calculus of variations. In other words, Jakobi established a link between the solutions of the original dynamics, the ODE is given by Hamilton's equation, and a new PDE for the action. PDE is standing for partial differential equation, of course, and ODE, ordinary differential equation. A link that forged out how we solve optimal control problems today. The Hamilton-Jakobi approach was essentially the 19th century precursor of what we now recognize as the dynamic programming of Hamilton-Jakobi-Belmann method. Hamilton also tackled the question of optimality more explicitly. After Euler and Lagrange found the stationary condition, it was natural to ask, is this extreme of a minimum or a maximum? As we mentioned, Legendre had already given an early test, requiring that the Lagrange and second derivative with respect to velocity be positive or negative, a kind of convexity condition. Jakobi went further by developing a more comprehensive criterion, involving what we know today is conjugate points, to ensure that an extremal path in the yields the minimum value of the action. In practical terms, Jakobi's work meant that one could check whether a proposed extremal path was really optimal or accidentally just a saddle point. Together, Hamilton and Jakobi forged a new calculus for physics, one in which dynamics, geometry and optimization were different phases of the same principle. Following Hamilton and Jakobi, the rest of the 19th century saw many further refinements. Notably, in the 1870s, Karl Weistras introduced his E-function and the so-called Weistras condition, which gave rigorous local tests for optimality of a variational trajectory. Weistras' condition is especially interested in hints at it. It essentially states that a long and optimal path, no infinitesimal variation of the trajectory, should decrease the Lagrangian, and viewed from an infinitesimal segment. In modern terms, this means that the integrand along the optimal path is locally maximum or minimal compared to the nearby paths. So Weistras formulated this in terms of this excess function, the E-function, but if one translates it into Hamilton's language, it says the following. For an optimal trajectory, at each instant, the chosen velocity or control makes the Hamiltonian as large or as small as possible. In fact, if we introduce a control input as a free choice of the velocity, Weistras' necessary condition can be refreshed as the optimal control maximizes the Hamiltonian at every point. And this is strikingly the same idea that many decades later would reappear at the heart of Pontragin's maximal principle. More on that later, of course. Yet, in the 19th century, the calculus of variations was still usually applied to systems where the control inputs were not explicit. One optimized overall possible curves satisfying endpoint conditions, but typically did not have any additional dynamically changing inputs bounded by constraints. The optimal choice of time-bearing controls as a separate problem was still waiting to be fully articulated. That stage was set in the 20th century when new technological challenges demanded a direct treatment of control variables. Fast forward to the mid-20th century. The space rays and guided missile areas demanded optimizing, really complicated dynamical systems. War and geopolitical rivalry now shaped research directions. In the Soviet Union, a mathematician named Lev Semyonovich Pontragin would emerge as a central figure in this context. Born in 1908, Pontragin lost his eyesight completely at the age of 14 after a tragic accident, the explosion of a gas stove followed by an unsuccessful eye surgery. But that didn't stop him. With extraordinary support from his mother, who taught herself mathematics just to help him inventing verbal descriptions of symbols so he could hear topology, Pontragin became one of the most accomplished mathematicians of his time. His mother dedicated herself to his education, as mentioned, and I think it would probably make sense to actually quote directly from the autobiography in Russian written by Pontragin himself, which is available by the way online and of course I'll provide a link in the description. Let me quote directly when Pontragin refers to the accident. From this moment Tatiana and Reyadna, his mother, assumed complete responsibility for ministering to the needs of her son in all aspects of his life. In spite of great difficulties with which she had to contend, she was so successful in herself appointed tasks that she truly deserves the gratitude of science throughout the world. For many years, in effect, she worked as a Pontragin's secretary, reading scientific works allowed to him, writing in the formulas in his manuscripts, correcting his work and so on. In order to do this, she had in particular to learn to read foreign languages. Tatiana and Reyadna helped Pontragin in all respects, seeing to his needs and taking very great care of him. Before left Pontragin revolutionized control theory, he had already traversed one of the most extraordinary paths in 20th century mathematics. In 1925, he entered Moscow University. He didn't take long for his professors to realize that he wasn't just good, he was exceptional. Imagine, a blind student who could not take notes, yet effortlessly retained pages of symbol heavy mathematics and not merely memorized them. He understood them deeply. His instructors were astonished not just as his ability to recall complex derivations, but on how clearly he grabs the underlying meaning behind the math. Leve first formed a deep intellectual bond with the mathematician Pavel Alexandrov, who would shape his early career. Alexandrov's charm, kindness and mathematical style focused on topology and geometric intuition left the last imprint. Pontragin thrived in this atmosphere, despite describing himself as less enthusiastic about kitchens and analysis for course. Alexandrov, meanwhile, found in Pontragin an ideal student, curious, rigorous and unafraid to wander into the frontier. By 1927, Pontragin then only 19 began making contributions to algebraic topology, specifically Alexandre duality. His key insight was from using linking numbers, a concept introduced by Broward, to understand the duality between homology groups of closed sets in Euclidean space and the homology of their compliments. After graduating in 1929, he was appointed to the Faculty of Mechanics and Mathematics at Moscow State. And just five years later, in 1934, he joined the Steklov Institute. We mentioned Steklov briefly in our episode on Lyapunov. Please go and check that out. That same year, he became head of its Department of Topology and Functional Analysis. His mathematical interests were broad, but his two passion lay in problems at the intersection of topology and algebra. Pontragin's work in this space wasn't just technically advanced, it was foundational. His duality theorems weren't mere exercises in abstraction, they enabled him to construct a general theory of character on locally compact, Abelian groups, paving the way for what been now called topological algebra. And this wasn't just an niche development. Historians now regard Pontragin's result as some of the most foundational mathematical advances of the century. To appreciate the scale of this, consider Elbert's fifth problem. One of the 23 famous problems Elbert posed in 1900. It asks whether every locally Euclidean topological group can be given the structure of a smoothly group. John Bonneumann cracked the compact group case in 1929, using his general integration theory. But it was Pontragin in 1934, who solved the problem for Abelian groups, using his new theory of characters. This was an astonishing leap forward, solving a Hilbert problem with an idea rooted in duality theory and harmonic analysis. Among his crowned achievements in this period was the 1938 publication of his book Topological Groups. It quickly became a classic, mathematicians across the globe praised it not just for its elegance, but for how it defined a new field. As one historian wrote in a quote, "This book belongs to that rare category of mathematical works that can truly be called classical books, that retain their significance for decades and exert a formative influence on the scientific outlook of entire generations." There's also a famous episode of Elie Cartan in 1934 visiting Moscow. Cartan gave a lecture in French, on computing the homology of classical compact league groups. Pontragin, who didn't speak French, said next to Nina Barry, who whispered a live translation. Cartan had proposed one approach, but Pontragin, still a young mathematician, returned the next year with a complete solution, using a totally different theory called Morse theory, in an entirely different way that Cartan had suggested. It wasn't just topology, where his name stuck. Pontragin's name is attached to multiple mathematical structures, Pontragin classes in differential topology, Pontragin duality in harmonic analysis, and the Pontragin-Tome construction. The conditional result in co-bordism theory. Some of these concepts would later be extended by mathematicians like Sergei Novikov, who proved the topological invariance of Pontragin classes, but Pontragin mostly laid the groundwork. So by the 1940s, Pontragin was a towering figure in mathematics, in particular in Soviet mathematics. But it was also closely connected to the Soviet state, a fact that would affect how his legacy was perceived the world, together with allegations of anti-Semitism. Lev rose through the ranks rapidly. In 1935, he was named the professor by the late 1940s, he was sharing editorial boards and leading research institutes. He became a member of the Academy of Sciences and was heavily involved in Soviet scientific policy. Then in 1952, Pontragin made a decision that somewhat stand many of his peers. He left pure mathematics to focus entirely on applied mathematics, especially differential equations and control theory. However, the shift had been brewing for decades. Since 1930s, Pontragin had collaborated with Alexandra Andronov, a pioneer in the theory of oscillations and nonlinear systems. The two had even co-authored the paper on dynamical systems as early as 1932. They maintained a close relationship frequently exchanging ideas and the boundary of physics and mathematics. But after Andronov's death, in 1952, Pontragin took up the torch more directly. That same year he began devoting his energies to become a mathematical study of automatic control systems. Back then, it was not necessarily a fashionable field for mathematicians, but it was a strategically important one. This wasn't just theory. The Soviet space program was already gearing up to launch Sputnik, which it did in 1957. Fuel-efficient rocket guidance wasn't a hypothetical question. It was an engineering priority. So the Soviet Union, now in throws of the Cold War, needed mathematical tools to control everything from aircraft to missile systems. Within a few years, it would produce the Pontragin Maximum Principle, the result that brought them into the world of optimal control and made him a household name. Not just among topologists, but among engineers, physicists and applying mathematicians around the world. So how did it all originate? There is a beautiful paper written by one of his closest collaborators, Rebats-Gam Krelitsse, called "The discovery of the Maximum Principle in Optimal Control". Of course, there will be a link in the description. But let me tell you a bit more about this story. So at this time of Institute in Moscow, Pontragin assembled a powerhouse group of collaborators, as mentioned, Rebats-Kamber-relitsse, but also Vladimir-Boliansky and Eugenie Mishenko among others. As mathematicians recount, Pontragin organized a seminar at the Institute where he regularly invited engineers and track practical problems. The activity in the seminar culminated in very soon in the formulation of two mathematical problems. One of them developed into the general theory of singularly perturbed systems of ordinary differential equations. The second problem brought the discovery of the Maximum Principle and the emergence of optimal control theory as we know it today. Let me quote directly now from this paper that I just mentioned, the discovery of the Maximum Principle in Optimal Control by Gam Krelitsse. Pontragin was led to the formulation of the general time optimal problem by an attempt to solve a concrete fifth-order system of ordinary differential equations with three control parameters related to optimal maneuvers of an aircraft. Which was proposed to him by two Air Force colonels during their visit to the Stakelov Institute in the early spring of 1955. Two of the control parameters entered the equations linearly and were bounded. Hence, from the beginning, it was clear that they could not be found by classical methods as solutions of the Euler equations. The problem was highly specific and very soon Pontragin realized that some general guidelines were needed in order to tackle the problem. I remember even said half jokingly, he quotes, "We must invent a new calculus of variations." As a result, the following general time optimal problem was formulated. And let me quote again because this is extremely interesting. Initially, it was supposed that the control back to U attains its value from an open set U. Necessity of a closed U, the most important case for applied problems, was evident from the beginning, though could be handled only later. To denote controlled parameters, the letter U was chosen as the first letter of the Russian world control, Uprovelinia. So what was this problem? Simply put, given initial and final states, x0 and x1 of a system, the goal was to find a control input U, possibly constrained to a given set, which drives the system from x0 to x1 in minimum time. Quoting again from Gamerellita's paper, the first and most important step towards the final solution was made by Pontragin right after the formulation of the problem. During three days, or better to say, during three consecutive sleepless nights, he suffered from severe insomnia, and very often he used to do math in bed all night. As a result, he completely disrupted his sleep in his later years and systematically took barbiturates in great quantities. Pontragin's maximum principle was born soon after. Pontragin and his team derived a set of necessary conditions for optimal control that generalized the Euler Lagrange condition. The key innovation was introducing adren variables, and for the technically oriented, this should be thought as the vector of Lagrange multipliers for the dynamics, and a Hamiltonian function for the control system. Pontragin showed that if U is an optimal control input driving the state x, then there must exist an adren vector P satisfying the adren differential equation, which simply states that the time derivative of the adren vector equals minus the derivative with respect to the state x of the Hamiltonian, along the optimal trajectory and the control. Which is analogous to Hamilton's canonical equations for the momentum. Furthermore, the optimal control input must maximize or minimize depending on convention the Hamiltonian and each monon time. In other words, the optimal control input must, at every moment, maximize this Hamiltonian. The system must evolve in such a way that the Hamiltonian is maximized with respect to the controls and every time step. Finally, certain transversality conditions must hold at the boundaries. For example, the moment the cos state variable or the adjoint variable at the final time must be orthogonal to the target set gradients. This was the first appearance of the adjoint system in optimal control theory, and it bridged optimal control to Hamiltonian formalism of classical mechanics explicitly. In fact, Prontiagin's method is often called an indirect method or a co-state method. You set up a Hamiltonian with state and co-state, or if you like, adjoint variables, and derive first order conditions for optimality. This concept of a co-state or adjoint state is now ubiquitous. For example, in economics where co-state variables are so-called shadow prices, and in machine learning where back propagation can be essentially seen as an adjoint method akin to the Pontiagin maximum principle. Mathematically, it's like using Lagrange multipliers, as mentioned earlier, for infinite-dimensional problems, with functions as variables extending Euler Lagrange equations by handling constraints, the system dynamics. The maximum principle conditions provide necessary conditions and are usually interpreted as a sort of generalized Euler Lagrange equations. They reduce to Euler Lagrange equations when applied to simple variational problems, but can handle control bounds and more general scenarios with ease. Pontiagin's maximum principle gives engineers and mathematicians a new powerful tool, a way to derive necessary conditions for an optimum in control problems with constraints. It doesn't always directly give the solution, but it reduces the infinite-dimensional optimization to solving a two-point boundary value problem, something similar to solving a differential equation that has a pre-specified initial state and a pre-specified final state. So now, instead of a single trajectory, one now has to solve a system of differential equations forward and backward in time, what became known as a shooting method. The mathematical formalism was rigorous and powerful, but the maximum principle also adds striking implications in practice. One famous consequence is the Bang Bang principle. If you want to minimize time or fuel, the best strategy is often to drive control inputs to their extremes, think full throttle or full brake, and never to casually cost. This explain, for instance, why a spacecraft might fire rockets at full power and then cost rather than hold an intermediate thrust. Go full throttle, then cut the engines, full burn or nothing. The principle justified were engineers at long-intuitive. One notable example is the Goddard rocket problem. Google that, you'll find it on Wikipedia, whereby one wants to maximize a rocket's altitude given limited fuel. Pontiagin's principle could derive the Bang Bang thrust schedule for that, something earlier methods struggle with. Pontiagin announced his maximum principle internationally in 1958 and the International Congress of Mathematicians in Edinburgh. The politics of the time made it tense. A blind Soviet mathematician talking about missiles at the Cold War peak. Western mathematicians at first greeted cooling, some dismissed it as just calculus of variations repackaged. Others may have been unsettled by Pontiagin's combative personality. He had a reputation for being brusque and outspoken, and by the suspicion that Soviet science by the 1950s was stated by secrecy and ideology. In truth, the situation was a bit more complex. By the late 1960s, his work appeared in the English language literature and his 1962 book, The Mathematical Theory of Optimal Processes, was translated by John Wiley in New York. As historian Bisson notes, Pontiagin's maximum principle was first published in the Open Soviet Literature in 1956, presented in English two years later and finally appeared in full proceedings in 1960. So once Western engineers saw what it could do, for example, in solving the Goddard rocket problem as mentioned earlier, or minimizing fuel for satellite trajectories, the power of the maximum principle became undeniable. And it is now regarded as one of the cornerstones of modern control theory. Notably, due to Cold War isolation, Western scientists were not initially aware of Pontiagin's breakthrough, and when they learned of it around 1960, some of them were skeptical. Naturally, but it quickly proved itself on various optimal trajectory problems. Meanwhile, it turned out that Harlem Hasteness in the US had independently developed similar ideas, the so-called calculus of variations with constraints. He had a form of the maximum principle as early as the 1950s, but his work didn't gain the fame that Pontiagin's did. By contrast, Pontiagin received the prestigious Lenin Prize for his work in 1961. Pontiagin's personal tale adds color to this development. As mentioned, a blind mathematician solving catynage engineering problems is remarkable in itself. Moreover, Pontiagin was known for his blunt and strong personality. At a 1970 conference after presenting work on differential games, he was asked by fame mathematician Alexander Grotendick about the morality of doing military research. Pontiagin retorted that if one were to follow that logic, and I quote, "we would be prohibited from speaking of abstract algebra either, since cryptography has military uses." Pontiagin saw the discovery of the maximum principle as an intellectual triumph of all, and indeed it was a huge triumph for Soviet mathematics, one that Soviet government publicized as evidence of leadership in the space-aged sciences. As we'll see next, almost at the same time, on the other side of the world, another approach to optimal control was being crafted, one that took a very different route via dynamic programming. While Pontiagin was developing his approach in Moscow, another revolutionary framework for optimal decision-making was taking shape across the Atlantic. In the United States, Richard Belman was crafting work would become known as the dynamic programming method, a method that would reshape how we approach time-dependent optimization even today. We explored Belman's life and legacy in a previous episode, please check it out, however it's worth revisiting dynamic programming here, because of just how profoundly his ideas impacted the theory and practice of optimal control. Belman's approach was fundamentally different, but ultimately complementary to Pontiagin's maximum principle. Belman's contribution is the so-called principle of optimality, the foundation, if you like, of dynamic programming, which he began formulating in the early 1950s. Belman's principle of optimality states, and now I quote, "an optimal policy has the property that whatever the initial state and initial decisions are, the remaining decisions must constitute an optimal policy with regard to the state resulting from the first decision." In other words, any prefix of an optimal trajectory is itself optimum for the subproblem starting at that prefix. This intuitive idea allowed Belman to break a complex multi-stage decision problem into simpler steps, a recursive decomposition. Instead of focusing on the instantaneous optimality conditions like Pontiagin's cost states at each moment, Belman focused on cumulative cost to go, from any given state. This insight leads directly to Belman's equation, a recursive functional equation for the optimal cost or value function, typically denoted as V. For a given state, Belman's equation says that the optimal value V is the minimum, over all possible control actions U, of the immediate cost, G of XU, plus the optimal future cost starting at the next state, which we may call F of XU. In other words, to act optimally now, you choose the action that minimizes the sum of what it costs you now and what it will cost you later, assuming you continue to act optimally from that point onward. Belman applied this to both discrete and continuous problems. In continuous control problems, his principle leads to the so-called Hamilton-Jakoby-Belman equation, we mentioned it earlier when we talked about Hamilton and Jacobi, and the HJB, as it's often called, is a partially financial equation for the value function, V of XT, the minimum cost from the state X at time T to reach the target. For example, in a time-optimal control, the value function V might represent the shortest time to reach the destination from the given state X at a time T. The Hamilton-Jakoby equation reads as follows. We have to take the partial derivative with respect to time of the value function and add to it a minimum, over all possible controls, of another function that, as mentioned, we when we call G of XU, which quantifies the cost, plus the partial derivative with respect to the state of the value function times the dynamics, say denoted by F of XU. All of this has to be equal to zero, with the terminal boundary condition that is given about the final cost. So V of X at capital T, capital T being the final time, has to be a certain function, F of X. Importantly, Belman's equation provides a necessary and sufficient condition for optimality under reasonable smoothness and convexity assumptions, unlike Pontraigin's condition, which are just necessary in general. However, the HJB equation must hold over the entire space, and this makes it notoriously difficult to solve, except in lower dimensions or special cases. In contrast, Pontraigin's maximum principle reduces to a boundary value problem along the optimal trajectory itself, which can be easier to solve and manage for specific trajectories. Therefore, the two approaches should be really regarded as complementary. Solving the Hamilton-Jacobie-Belman equation gives you a global solution and a closed loop optimal feedback law. However, solving the Pontraigin maximum principle equations, give you an open loop optimal trajectory, which can be converted to feedback if needed via the value function you obtain on that trajectory. In many classical settings, especially for smooth, finite horizon control problems with convex costs and dynamics, they yield equivalent necessary conditions for optimality. For instance, solving the Hamilton-Jacobie-Belman equation for well-behaved linear quadratic problems yield the same ricati equation or optimal feedback law that emerges from Pontraigin's two-point boundary value problem. When Belman first proposed dynamic programming around 1953, summing the mathematical community where initially doubtful, it looked less rigorous than calculus of variations, Belman was essentially guessing a value function existing and then writing an equation for it. But he championed it fervently and applying dynamic programming to basically everything from planning to inventory management. An amusing anecdote from Belman's autobiography, which we cover again in the previous episode of "In Control Dedicated to Belman", he describes how he coined the term "dynamic programming". The idea was to at least impart secure funding and avoid political pushback. In the 1950s, the word programming meant devising plans as in military programs and dynamic at a flashy positive ring. Belman reiterated that a US Secretary of Defense had time hated the term and a "research". So, Belman deliberately chose an opaque name for his work to make it more palatable, quoting again from his autobiography and provide a link in the description, one also gets a sense of the importance of choosing names. What title, what name could I choose? In the first place, I was interested in planning, in decision-making, in thinking. But planning is not a good word for various reasons. I decided therefore to use the word programming. I wanted to get across the idea that this was dynamic, this was multi-stage, this was time-bearing. I thought, let's kill two birds with one stone. Let's take a word that has an absolutely precise meaning, namely dynamic in the classical physical sense. And he also has the very interesting property as an objective that it's impossible to use the word dynamic in a pejorative sense. Try thinking of some combination that will possibly give it a pejorative meaning. It's impossible. Thus, I thought dynamic programming was a good name. It was something not even a congressman could object to. So I used it as an umbrella for my activities. In 1957, Richard Belman published his landmark book, Dynamic Programming, alongside a flurry of papers that quickly seeded his ideas across disciplines. His recursive method for solving sequential decision problems found fertile ground in economics, operation research and of course control engineering. But Belman also issued a prescient warning. The now famous curse of dimensionality, as the numbers of the state variable increases, the computational burden of dynamic programming grows exponentially. That curse still haunts us today and remains a central obstacle in high-dimensional control problems. Despite this, dynamic programming proved remarkably effective in lower-dimensional systems and in cases where clever approximations could be brought to bear. As we shall see later, Belman's ideas would experience a stunning 21st century revival under the banner of reinforcement learning, which at its core amounts to solving Belman's equation using simulation and statistical approximation rather than symbolic computation. In the decades that followed, researchers essentially had two main roads, solved contraagans 2.1 boundary value problem or Belman's recursive value equation, depending on the nature of the system and the tools at hand. Or, if you were Rudolph Kalman, as we shall see next, you found a way to bridge both worlds, solving them simultaneously in a special, elegant case. No history of optimal control is complete without mentioning Rudolph Kalman, who in 1960 at the age 29 published two landmark works that tied everything together and propelled control theory in the modern era. Kalman's first paper, Contribution to the Theory of Optimal Control in 1960, addressed the linear quadratic regulator or briefly Alqr problem, using state space methods. His second, a new approach to linear filtering and prediction problems in 1960 and then later in 1961 with Busy, introduced the Kalman filter or the Kalman Busy filter for optimal state estimation. Together, this created the linear quadratic Gaussian or briefly Alqg framework that still is an important pillar of modern control theory today. The Kalman filter algorithm provided an optimal recursive solution for estimating the state of a linear dynamic system from noisy measurements. While, strictly speaking, the Kalman filter is about estimation, a dual problem to control, it's rooted in optimal control reasoning, specifically it solves a least squares minimization problem, an optimal estimation problem, which is the dual of an Alqr control problem. The filter was revolutionary for navigation and guidance systems. NASA famously quickly adopted the Kalman filter in the Apollo program for the navigation of the Apollo spacecraft. It allowed the onboard computer to continuously estimate the capsule's trajectory and make corrections, using data from multiple sensors. This is a direct real-world payoff of optimal control theory. Without Kalman filters and the optimal control theory behind them, navigating to the moon with the limited computing of the 1960s would have been far less accurate. If you're curious, we have died into some of these developments in episode 34, please go and check it out. Around the same time, Kalman tackled, as I mentioned, another challenge in optimal control, the Alqr problem. This involves designing a feedback controller that minimizes a quadratic cost, typically a weighted combination of state deviation and control effort, for a linear dynamical system. In 1960, Kalman showed that this problem admits a remarkably elegant solution. The optimal control law is linear in the state, and it can be always written in terms of a function that depends on a matrix-valued time function P, that satisfies a differential Riccati equation, or an algebraic Riccati equation at steady state. This wasn't just a clever trick. It unified and formalized what had previously been heuristic design practices. The derivation could be carried out using both Balmons dynamic programming framework or Pontraagin's maximum principle, and both led naturally to the same Riccati equations. Control engineers now had something that they hadn't had before, a systematic way to design feedback controllers that were provably optimum. To appreciate this significance, consider the context. Before the 1960s, control design in aerospace, electronics, and industry was largely the domain of classical control. Frequency domain methods like body plots and the Nyquist criterion. These were powerful tools, but largely unhawk and focus primarily on stability. Kalman, in contrast, championed a new approach, what became known as "modern control theory". He advocated for working in state space domain, directly modeling systems via equations like x dot equals a x plus b u, and using linear algebra to understand controllability of the system. Controllability, observability, and optimality. Concepts he introduced and formalized in the 60s. By the mid-late 60s, the tide had turned. Optimal control began showing up in aerospace curricula and textbooks. The IEEE launched the control systems magazine to highlight these emerging methods, and the field began to organize around this new paradigm. Of course, Kalman wasn't alone in this transformation. Researchers like John Ragazini and Lofi Zadeh had laid important groundwork in sample data systems and early state space formulations. But Kalman's results provided a mathematical clarity and computational tools that solidified the field. In 1962, he and Richard Busy published the Continuous Time Kalman Busy Filter. A few years later, Bryson and Ho released their influential 1969 textbook "Applied Optimal Control", which brought tools like the Pontragin Maximum Principle and Ricati-based methods into the hands of practicing engineers. Researchers like Peter Fald and Mike Athens, as well as Brian Anderson and others, followed developing fast numerical algorithms to actually compute Ricati solutions and deploy them in real time. After the foundational breakthroughs of the 60s, the field of optimal control expanded to tackle the messy realities of real world systems. Physical systems are often linear, operate under constraints on both states and controls, and frequently involve non-smooth behavior, for example, far from the idealized quadratic or linear models of early optimal control. So from the 17th onwards, all the way to today, a wave of mathematicians, physicists, and engineers help extend the theoretical machinery to these broader, more challenging settings. So in what follows, I'm going to mention a few, a necessarily limited number of key advances related to optimal control. Let's start with Rufus Isaacs and differential games. Although his work began in the 50s, Rufus Isaacs pioneering theory of differential game did not receive widespread recognition until later decades. Like Batman, Isaacs was an mathematician at Rand Corporation, and he was tackling the new kind of problem. How do you formulate optimal strategies when there are two intelligent agents with conflicting goals? His 1965 book, "Differential Games" became the foundational text in this area, introducing the so-called minimax principle essentially a two-player version of Belman's dynamic programming, and the Hamilton-Jakomi-Belman equation. In this setting, the value function satisfies the Hamilton-Jakobi Isaacs equation, a partial differential equation encoding the optimal strategy for each player in a zero-sum game. His approach was novel, Isaacs re-framed optimal control as a game against an adversary. Instead of assuming a single agent trying to minimize cost in a known environment, differential games allowed for competing agents, one trying to minimize costs and the other trying to maximize it. This naturally modeled pursuit evasion problems, military engagements, and even robust control where the opponent may be interpreted as the worst-case disturbance or uncertainty in the system. Isaacs also introduced several key mathematical ideas such as barriers, programmed iteration and strategies under uncertainty, which led to the groundwork for future developments in both robust control theory and game-fieretic planning. His work inspired whole lines of research in economics where markets are modeled as games, aerospace, where missile interception, and tracking can be posed as differential game problems, and robotics, whereby evasive maneuvers, multi-agent systems, and adversarial learning lend themselves very naturally to game-theoretic modeling. I'm going to mention now another figure that is probably more well-known in the mathematics field, but definitely deserves recognition also from the control community. In the 1970s, Francis, also known as Frank Clarke, fundamentally reshaped the mathematical landscape of optimal control with the creation of non-smooth analysis, a new calculus for a less idealized world. Classical optimal control theory from Hamilton to Pontreagge into Bellman largely assumed smoothness, that cost functions, system dynamics, and value functions were differentiable at least. But many real-world problems violate this assumption. Control systems often face hard constraints, on states, on controls, on time, and can involve discontinuous actions, like switching or saturation, for example. These resulting value functions frequently have kinks, casps or flat regions, and standard tools like gradients simply break down. Clarke addressed this with a profound insight. Even if a function isn't differentiable in the classical sense, it still possesses meaningful structure. Introduced a notion of generalized gradient, now widely known as the Clarke sub-differential, to describe the slope of non-smooth functions. This opened the door to a new kind of calculus, one that could handle optimization in settings where the classical derivative doesn't exist. Clarke didn't stop at analysis, of course. He applied these tools to extend Pontreagge's maximum principle to far more general classes of systems. Including those described by differential inclusions where the control doesn't appear as a unique function, but rather as a set-valued map. This was essential for handling hybrid and discontinuous systems where the evolution of the state can switch abruptly due to mode changes, threshold events, or logic-based constraints. In practical terms, Clarke's contributions allowed researchers to derive necessary conditions for optimality, even in problems with non-smooth costs. And constraints. His landmark 1983 book, Optimization and Non-Smooth Analysis, remains a foundational text in the field. It laid out the theoretical machinery including sub-differentials, normal cones, and non-smooth Euler Lagrange conditions that has since become standard in model-contron theory and generated the wealth of research in optimization. Clarke later authored a comprehensive treatise on optimal control that further codified these ideas and extended them to dynamic systems, partial differential equations, and constrained control. Another significant development that we definitely need to mention is, of course, model predictive control, which we discussed at length in an earlier episode with Manfred Morari. Model predictive control or MPC, as we call it today, has its roots in the 70s and 80s process industries, oil refineries, but also chemical plants, where engineers faced multi-variable optimal control problems with constraints like inputs and state payments. Early versions of MPC like dynamic matrix control, also called DMC by Cutler and Remaker at Shell, in the 70s where essentially heuristic optimal controllers. Every few seconds they solved an optimization problem, using typically a linear model aquatic cost over a finite horizon to compute the next control action and then roll the horizon forward. By the late 80s MPC was essentially formalized, you choose a prediction horizon set up a cost function, for example tracking error plus input effort, and solve a constrained optimization problem, which is typically quadratic at each time step. The result is an optimal control sequence, but you apply it only for the first input and then you re-optimize at the next step. This provides feedback lows while handling constraints gracefully. Initially people worried we have to solve an optimization problem in real time, this is not really possible. But thanks to Moore's law by the 90s, even complex plants could run MPC controllers in milliseconds. Academics like Manfred Morari, as we mentioned, but also David Main and John Rollings, help put MPC on a firm theoretical footing. For example, providing conditions to ensure stability. MPC in essence applies optimal control in a receding horizon fashion. It doesn't find the infinitorizen optimal policy in one go, it repeatedly solves finite horizon approximations of such a problem. Amazingly, in many applications, this works extremely well, and by the 2000s MPC became essentially the go-to method for everything from controlling autonomous vehicles to managing power grids. Historical note is also the following, the term MPC came later. The early industrial names were model predictive, heuristic control, dynamic matrix control, and many others. But the concept illustrated how optimal control ideas left the lab and entered daily use in factories worldwide. Today's self-driving cars, for instance, often use MPC to compute steering and acceleration commands that minimize trajectory tracking while respecting physical constraints. Finally, we have to talk about the relationship between optimal control and reinforcement learning. One of the most transformative developments in recent decades, and perhaps the most striking bridge between optimal control and computer science, is the rise of reinforcement learning, or RL. At its art, RL is just optimal control, but without a model. Rather than relying on system dynamics, one learns behavior through trial and error. If Belman's dynamic programming laid theoretical foundation, RL is its modern data-driven hair, pushing Belman's ideas into the age of algorithms and neural networks. In RL, an agent learns to interact with an environment, a system that may be stochastic or only partially known, by attempting to maximize a cumulative reward over time. The Belman equation sits at the very core of this process, providing the recursive logic behind almost every RL method. To learning, for instance, explicitly seeks to find the optimal Q function that satisfies the Belman equation. Even more advanced, methods like policy gradient or actor-critic methods rely on approximating value functions, or advantage estimates that trace back to Belman's formulation. Monte Carlo III's research, for example, used famously in DeepMind's AlphaGo, is in many ways a form of empirical dynamic programming. What's remarkable is how this blend of learning and control has conquered problems once thought untouchable. In 2016, DeepMind's AlphaGo, if you don't know it, of course I'll provide a link in the description that's a very nice YouTube movie about it, AlphaGo shocked the world by defeating the World Champion at the ancient board game of Go, and the main with more legal positions than atoms in the universe. At its core, AlphaGo combined deep neural networks with value function estimation and simulation-based planning, casting the Go board as a high-dimensional state space and every move as an action, making it essentially an optimum control problem disguised as a game. But the impact of RL doesn't stop at games. Today, in robotics, RLs.legged machines out of walk and jump, drones out to perform agile maneuvers, and cars out to navigate complex roads, all by maximizing long-term objectives like energy efficiency, safety, and task completion. In science and engineering, RL is being used now to try and control plasma in nuclear fusion reactors, optimized molecular designs, and scheduled tasks in cloud computing centers. Even in operations research, a field with rep roots in control theory, RL re-emerged under the name of approximate dynamic programming, helping tackle resource allocation problems where classical solutions buckle under the weight of complexity. Crucially, RL didn't just borrow from optimal control. It's feeding ideas back into the field. Concepts like exploration versus exploitation, adaptive learning, and dual control where a system must learn and act optimally at the same time are being revisited through the lens of RL. Neural networks, once with suspicion by control theorists, are now being used to approximate value functions and policies, offering an medical way to solve Hamilton-Jacobie-Bellmann equations in problems that defy classical approaches, for example humanoid robotics. In this way, the story seems to have come full circle. Re-emforcement learning may seem new by its mathematical DNA is pure optimal control. From Bellmann's recursion to modern D. Parallel, the threads remain. How do we act optimally over time, especially when the world is uncertain? Thanks to modern computation, RL has breathed new life into that ancient question and brought optimal control theory to the very heart of the AI revolution that we're living through today. So, from Johann Bernoulli-Brakistokron challenge in 1697 to two days, reinforcement learning agents and biologically inspired control systems, we've traced a remarkable evolution. Optimal controls history is a testament to human curiosity and ingenuity. It began in the age of enlightenment, with mathematicians daring each other to find optimal paths and shapes, and in the process inventing a new kind of mathematics, it matured through the industrial revolution, as figures like Hamilton and Jacobi wove it into the fabric of classical physics, showing that nature itself seems to evolve along optimal trajectories. It left forward through the Cold War, when pioneers like Pontreaggin and Bellmann working continent apart formalized their rigorous theory of decision making over time for everything from rocket guidance to economic planning. It was made practical by people like Kalman and others who brought a modern computational mindset, merging abstract theory with the tools needed to actually implement control, and in recent decades optimal control has spread far beyond its aerospace roots, influencing medicine, neuroscience, finance, and artificial intelligence. From the brain of a fruit flying to the moves of AlphaGo, the same core questions how to act optimally over time, at playing out in surprising places. But this story, as rich as it is, still leaves a lot untold. We didn't even touch the origins and significance of the Euricati equation. In the deep connections between optimal control and dissipativity theory, a concept that links energy, stability, and optimality. We didn't even explore the intriguing world of inverse optimal control, where the goal is to infer the cost function and agent is optimizing, a crucial idea in fields like neuroscience and imitation learning. Nor did we dive into the many modern extensions of optimal control, stochastic control, hybrid systems, non-linear feedback design, and the frontier of control under information constraints. So now, I'll turn the mic to you. What did we miss? What part of this story deserves more light, more time, or even a whole episode of its own? Let me know. When if you enjoyed this journey, please give us a like, a follow, or a five-star review whenever you're listening. It helps more curious minds discover the in-control podcast. Until next time. Thank you for listening. I hope you liked this show today. If you enjoy the podcast, please consider giving us five stars on Apple Podcasts, follow us on Spotify, support on Patreon, or Paypal. And connect with us on social media platforms. See you next time.
Podcast Summary
Key Points:
The podcast introduces optimal control theory as the discipline addressing how to act optimally over time, with roots tracing back over 300 years.
The 1696 Brachistochrone problem, posed by Johann Bernoulli, is highlighted as a foundational moment, marking the shift from static geometry to dynamic, time-optimizing problems under physical constraints.
The episode explores the historical development through key figures like the Bernoulli family, Newton, Leibniz, Euler, and Lagrange, noting their rivalries and contributions that shaped the calculus of variations and modern optimal control.
It discusses the cultural and academic conflicts of the era, particularly the calculus priority dispute between Newton and Leibniz, and how these influenced mathematical progress.
The narrative connects early problems like the Brachistochrone to later 20th-century breakthroughs such as Pontryagin's maximum principle and dynamic programming, framing them as a continuous search for optimality conditions.
Summary:
This podcast episode traces the origins and evolution of optimal control theory, beginning with its philosophical core: determining the best actions over time. It identifies the 1696 Brachistochrone problem—finding the curve of fastest descent under gravity—posed by Johann Bernoulli as a pivotal moment. This challenge moved beyond static geometry to optimize a dynamic process (time) under physical laws, effectively birthing the field.
The story unfolds through historical narratives involving key mathematicians like the Bernoulli family, Newton, Leibniz, Euler, and Lagrange, highlighting their intellectual rivalries, such as the bitter calculus priority dispute between Newton and Leibniz. The episode explains how these early efforts in the calculus of variations systematically developed methods for optimizing functionals, setting the stage for 20th-century revolutions like Pontryagin's maximum principle and dynamic programming. It frames the entire history as a sustained quest to formalize necessary conditions for optimality, connecting a 17th-century puzzle directly to the mathematical foundations of modern engineering, physics, and computation.
FAQs
Optimal control theory is the discipline of making decisions over time to achieve the best possible outcome, with roots stretching back over 300 years. It provides mathematical foundations for developments in physics, engineering, and computation.
The brachistochrone problem, posed by Johann Bernoulli in 1696, asks for the curve of fastest descent between two points under gravity. It is considered the birth of optimal control because it framed optimization in terms of time and dynamic behavior under physical laws.
Key figures include Johann Bernoulli, who posed the brachistochrone problem; Leonhard Euler and Joseph-Louis Lagrange, who systematized the calculus of variations; and later contributors like Pontryagin, Bellman, and Kalman in the 20th century.
The solution to the brachistochrone problem is the cycloid curve. This curve also has the property of being a tautochrone, where oscillation period under gravity is independent of starting point, showing nature's tendency toward simplicity.
The rivalry over calculus invention led to disputes, with Johann Bernoulli supporting Leibniz. This tension partly motivated Bernoulli's brachistochrone challenge, aiming to showcase Leibniz's calculus and provoke Newton, who anonymously solved it.
The Bernoulli family produced eight notable mathematicians across three generations, including Jacob, Johann, and Daniel Bernoulli. Johann's brachistochrone challenge initiated optimal control, while their work in calculus, probability, and fluid mechanics was foundational.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.