---
title: 'Advanced: optimal control and MPC'
url: https://doc.liz6.com/en/theory/02-control-theory/11-optimal-and-predictive-control
locale: en
area: theory
tags:
- Control theory
- Theory
date: 2026-09-10
modified: 2026-09-10
description: Pole placement specifies dynamics, but not the relative cost of error and action. Optimal control makes that tradeoff explicit; MPC includes a finite future and constraints in repeated decisions. This chapter builds on state feedback, digital control and quadratic functions.
---

# Advanced: optimal control and MPC

Pole placement specifies dynamics, but not the relative cost of error and action. Optimal control makes that tradeoff explicit; MPC includes a finite future and constraints in repeated decisions. This chapter builds on [state feedback](10-state-feedback-and-observers.md), [digital control](07-sampling-and-delay.md) and quadratic functions.

## LQR: define the objective

For continuous $\dot x=Ax+Bu$, consider

$$J=\int_0^\infty(x^\mathsf TQx+u^\mathsf TRu)dt,\qquad Q\succeq0,\ R\succ0.$$

$Q$ penalizes state deviation and $R$ action. Coordinates and units matter: changing meters to centimeters without changing weights changes the physical preference. Allowable errors and actions can provide normalization scales.

Under standard stabilizability and detectability conditions, the stabilizing Riccati solution $X$ satisfies

$$A^\mathsf TX+XA-XBR^{-1}B^\mathsf TX+Q=0.$$

The feedback is $u=-Kx$, $K=R^{-1}B^\mathsf TX$. Here $X$ is a cost matrix, not the state. Basic LQR has no hard actuator constraints; clipping its result changes its optimality and guarantees.

## A hand-computable optimum

For $\dot x=u$, $Q=q>0$, $R=\rho>0$, the equation is $-X^2/\rho+q=0$. The stabilizing solution gives

$$X=\sqrt{q\rho},\qquad u=-\sqrt{q/\rho}\,x.$$

At $q=\rho=1$, $x=e^{-t}x_0$. Raising $\rho$ to 4 reduces gain to 0.5 and slows convergence. For $x_0=2$, the respective optimal costs are 4 and 8. Their objectives differ, so those numbers cannot directly rank the controllers under one common preference.

## MPC: optimize a sequence, execute one action

For discrete $x_{k+1}=Ax_k+Bu_k$, solve at each instant

$$\min_{u_{0:N-1}}\sum_{i=0}^{N-1}(x_i^\mathsf TQx_i+u_i^\mathsf TRu_i)+V_f(x_N)$$

subject to $x_0=\hat x_k$, dynamics, $x_i\in\mathcal X$, $u_i\in\mathcal U$, and any terminal constraint. Execute only the first action, then measure or estimate again and reoptimize.

This incorporates new information after disturbances or prediction errors and can use known future references. Predicted and realized trajectories are different objects. The [do-mpc introduction](https://www.do-mpc.com/en/latest/theory_mpc.html) describes finite horizons, estimation and constraints.

## Inspect a finite-action example

Our dimensionless model is $x_{k+1}=x_k+u_k$, $x_0=2.5$, target 0, with actions $\{-1,-0.5,0,0.5,1\}$. Define the specific cost

$$J_N=\sum_{i=0}^{N-1}\left(x_{i+1}^2+\rho u_i^2\right).$$

There is no extra terminal term. Unlike the preceding notation, this penalizes each post-action state. Searching all $5^N$ sequences finds the global optimum within this finite set, not the general continuous-input MPC optimum.

**Advanced: optimal control and MPC · Experiment**

x(k+1)=x(k)+u(k), x0=2.5; actions 0, ±0.5, ±1; cost Σ[x(i+1)²+ρu(i)²]. Search all 5^N candidates, execute the first action, and reoptimize. Dashed initial prediction has no advance knowledge of the disturbance.


Change horizon and action weight. Compare the initial prediction with the realized receding-horizon trajectory and actions. An optional +1 disturbance is applied after the third action; it is not predicted. At equal cost, prefer smaller absolute first action, then enumeration order.

For $N=1,\rho=1,x=2.5$, action −1 costs $1.5^2+1=3.25$, while −0.5 costs $2^2+0.25=4.25$. Choose −1. For continuous input constrained to $[-1,1]$, the stationary point is $u=-x/(1+\rho)$, projected onto the interval. A finite set additionally requires comparing neighboring candidates.

## Feasibility and stability need their own arguments

A feasible optimization now does not ensure feasibility after a disturbance: recursive feasibility is a separate property. Finite-horizon optimality also does not automatically imply infinite-time convergence. Suitable terminal costs, terminal sets and local controllers are common ingredients in guarantees, with model-specific conditions.

Longer horizons see farther but cost computation. Missing the sample deadline introduces delay. Define fallback actions for timeout or infeasibility, constraint priorities, estimation uncertainty and model updates. Soft constraints need explicit allowed violations and penalties; do not silently soften hard limits.

## Check your understanding

Is a clipped LQR action still the optimum of the original unconstrained problem?

<details><summary>Reasoning</summary>

No. Clipping changes the loop. Analyze saturation or formulate the constrained problem, then reassess stability and feasibility.

</details>

Why not compute one sequence and execute it forever?

<details><summary>Reasoning</summary>

Disturbances, model errors and updated estimates change the state. Repeated measurement and optimization use that new information; computation delay remains part of the implementation.

</details>

## Reproduce and continue

[python-control `lqr`](https://python-control.readthedocs.io/en/stable/generated/control.lqr.html) returns gain, Riccati solution and closed-loop eigenvalues. This checks the scalar example using SciPy:

```python
import control as ct
K, X, poles = ct.lqr([[0.0]], [[1.0]], [[1.0]], [[4.0]], method="scipy")
print(K, X, poles)
# [[0.5]] [[2.0]] [-0.5]
```

Continue with system identification for better models, nonlinear/robust control for broader guarantees, or stochastic estimation for noise. Return to the [reading route](index.md) to choose a path.
