Getting Started With Optimal Control Theory

I spent about three years working on trajectory optimization for satellite maneuvers before I actually understood what was going on. The math looks clean on paper but hits a wall pretty quickly when you try to run it on real hardware. So here is how I approached Optimal Control Theory An Introduction, what tripped me up, and what actually works. Optimal control theory is fundamentally about finding the best possible input for a dynamic system over time. You have a set of differential equations describing how the system moves. You pick a cost function that penalizes things you do not want. Then you minimize that cost subject to the dynamics and any constraints you might have on states or inputs. That is it in the broadest sense. The difficulty comes from the details.

Optimal Control Theory An Introduction

There are two main families of approaches you will run into. The direct methods and the indirect methods. Each has serious trade-offs that beginners tend to gloss over. Direct methods convert the continuous problem into a finite-dimensional nonlinear program. You parameterize the control and maybe the state, discretize the dynamics using something like collocation or shooting, then hand it to an NLP solver. I usually reach for GPOPS-II or ACADO for this because they handle the transcription nicely and interface with IPOPT without much fuss. The downside is that direct methods can be slow to converge on hard problems, and you have no guarantee you found the global optimum. You just found something the solver liked. Indirect methods go the other direction. You derive the necessary conditions using Pontryagin's minimum principle, end up with a two-point boundary value problem, and solve that directly. The benefit is you get analytic structure and typically faster convergence when you are close to the solution. The problem is deriving those conditions correctly. One sign error in the Hamiltonian and your costate equations are garbage. I have seen people spend weeks debugging indirect formulations only to find a single misplaced negative sign.

The Hamiltonian itself is H = L + ^T f, where L is your running cost, f is your dynamics, and is your costate. The optimality condition says the control minimizes H pointwise. The costate evolves according to the negative partial derivative of H with respect to the state. Boundary conditions come from transversality conditions at the endpoints. This is standard textbook stuff, but applying it to anything beyond textbook examples requires care. Here is something most tutorials skip: constraint qualification. When you have path constraints like q(u,t) 0, you need to handle them properly. A common mistake is to just slap on a penalty term and hope for the best. That rarely works well. Instead, you introduce additional Lagrange multipliers and deal with the resulting complementarity conditions. Slightly more work, but you actually respect the constraints. I ran into a specific issue once while optimizing a spacecraft rendezvous trajectory with thrust magnitude bounds. The problem had a singular arc where the Hamiltonian did not depend explicitly on the control direction. Standard software kept oscillating because it could not decide which way to point the thrust. I ended up using a regularization technique where I added a small quadratic term to the control in the singular region, then gradually reduced it. Took about four hours to implement but solved the whole problem cleanly. There is also the option of using higher-order generalized Legendre-Clebsch conditions to detect singular arcs analytically, but that requires deriving third derivatives of the Hamiltonian, which is painful.

Get the Full Details

Optimal Control Theory: An Introduction (Dover Books on Electrical Engineering) : Kirk, Donald E ...
Optimal Control Theory: An Introduction (Dover Books on Electrical Engineering) : Kirk, Donald E ...

Another thing to watch out for is numerical scaling. Optimal control problems often have variables with wildly different magnitudes. Position might be in kilometers, velocity in meters per second, thrust in kilonewtons, and time in seconds. Solvers struggle with this. I usually normalize everything to be roughly order one before handing the problem off. It takes five minutes and prevents a lot of headaches later. When you do need to handle uncertainty, stochastic optimal control adds a layer of complexity. You are no longer minimizing a deterministic cost. You are minimizing an expected cost, which usually means resorting to dynamic programming or Monte Carlo sampling. Both have issues. Dynamic programming suffers from the curse of dimensionality. If your state space has more than about ten dimensions, you are likely out of luck unless you use approximations like Q-learning or policy gradients. Those trade off guaranteed convergence for practical applicability. Receding horizon control, or MPC, is a pragmatic workaround for many real-world applications. You solve an open-loop optimal control problem over a finite horizon, apply the first chunk of the control, then repeat at the next time step. It is suboptimal by design but handles constraints and disturbances much better than open-loop optimization. I used MPC extensively for quadrotor trajectory tracking. The computational cost is nontrivial. Even with warm-starting and efficient NLP solvers, you are looking at 50 to 200 milliseconds per solve on modern hardware, depending on horizon length and problem size.

If you want to learn this stuff properly, start with Bryson and Ho for the classical treatment. It is old but thorough. For a modern computational perspective, try Betts' book on direct methods. It covers transcription schemes in detail and includes practical implementation notes. Online, the ICLOCS workshop proceedings are a good source for recent advances. The ISOPT conference also has useful material. Software-wise, if you are just getting started, try PROPT in MATLAB. It is built on top of GPOPS and has a decent user interface. For Python, the scikits.optimization package is okay but somewhat limited. CasADi is probably the most flexible option, though the learning curve is steep. If you are willing to invest time in C++, IBIP and KINSOL give you more control over the solver internals. The hardest part of optimal control is not the theory. It is making the numerics work. Convergence issues, constraint violations, poor initial guesses, bad scaling. These will bite you repeatedly. My advice is to start simple. Solve a double integrator problem with fixed endpoints. Verify your code against the analytic solution. Then gradually add complexity: dynamics, constraints, terminal costs. Do not jump straight into a six-degree-of-freedom vehicle model and expect things to work.

Also, keep good notes on your initial guesses. A poor guess can make or break convergence, and you will forget why you chose it. Document your parameters, your discretization choices, and any regularization you applied. Three months from now, you will thank yourself.

Optimal Control Theory: An Introduction by Donald E. Kirk
Optimal Control Theory: An Introduction by Donald E. Kirk