0% found this document useful (0 votes)
1 views364 pages

Lecture Notes On Stochastic Opt Control

These lecture notes cover the optimization and control of stochastic systems, including probability theory, controlled Markov chains, and classification of Markov chains. The document is structured into sections that review essential concepts, introduce controlled models, and discuss performance criteria. It also includes exercises and bibliographic notes for further study.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
1 views364 pages

Lecture Notes On Stochastic Opt Control

These lecture notes cover the optimization and control of stochastic systems, including probability theory, controlled Markov chains, and classification of Markov chains. The document is structured into sections that review essential concepts, introduce controlled models, and discuss performance criteria. It also includes exercises and bibliographic notes for further study.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Optimization and Control of Stochastic Systems

Serdar Yüksel
Queen’s University, Mathematics and Statistics

Lecture Notes

April 29, 2026


ii

These notes have been been prepared for MTHE/MATH 472 / MATH 872: Optimization and Control of Stochastic Systems at
Queen’s University and also used for EEE 446/546: Control and Optimization of Stochastic Systems at Bilkent University.
Contents

1 Review of Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 Measures and Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2.1 Borel σ-field . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.2.2 Measurable Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.2.3 Measure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.2.4 The Extension Theorem (Optional) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.2.5 Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.2.6 Fatou’s Lemma, the Monotone Convergence Theorem and the Dominated Convergence Theorem . . 7
1.3 Probability Space and Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
1.3.1 More on Random Variables and Probability Density Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
1.3.2 Independence and Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
1.4 Stochastic Processes and Markov Chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
1.4.1 Markov Chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
1.5 Appendix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.5.1 Proof of Theorem 1.2.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.6 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11

2 Controlled Markov Chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15


2.1 Controlled Markov Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.2 Fully Observed Markov Control Problem Model (MDP Models) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.2.1 Classes of Control Policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
2.3 Performance Criteria: Optimality and Stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
2.3.1 Several Optimality Criteria and Performance of Policy Classes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
2.3.2 Stability as a Performance Criterion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
iv Contents

2.3.3 Markov Chain Induced by a Markov Policy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18


2.4 Partially Observed Models and Reduction to a Fully Observed Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
2.5 Decentralized Stochastic Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.6 Controlled Continuous-Time Stochastic Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.7 Numerical Methods, Reinforcement Learning, and Robustness to Incorrect Models . . . . . . . . . . . . . . . . . . . . 20
2.8 Bibliographic Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.9 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20

3 Classification of Markov Chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25


3.1 Countable State Space Markov Chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
3.1.1 Recurrence and transience . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
3.1.2 Stability and invariant measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.1.3 Invariant measures via an occupational characterization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.1.4 Rates of convergence to invariant measures and Dobrushin’s ergodic coefficient . . . . . . . . . . . . . . . . . 31
3.1.5 Ergodic theorem for countable state space chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
3.2 Uncountable Standard Borel State Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
3.2.1 Invariant probability measures and split chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
3.2.2 Existence of an invariant probability measure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
3.2.3 On small and petite sets: sufficient conditions (Optional) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
3.2.4 Rates of convergence to equilibrium . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
3.3 Further Conditions on the Existence and Uniqueness of Invariant Probability Measures . . . . . . . . . . . . . . . . . 42
3.3.1 Further conditions on existence of invariant probability measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
3.3.2 Uniqueness of an invariant probability measure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
3.4 Ergodic Theorems for Markov Chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
3.4.1 Ergodic theorems for positive Harris recurrent chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
3.4.2 Further ergodic theorems for Markov chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
3.5 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48

4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics 51
4.1 Martingales . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
4.1.1 More on expectations and conditional probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
4.1.2 Some properties of conditional expectation: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
4.1.3 Discrete-time martingales . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54
4.1.4 Doob’s optional sampling theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
4.1.5 Doob’s maximal inequality (optional) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
Contents v

4.1.6 An important martingale convergence theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57


4.1.7 The ergodic theorem[Optional] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
4.1.8 Further martingale theorems [Optional] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
4.2 Stability of Markov Chains: Foster-Lyapunov Techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
4.2.1 Criterion for invariance (existence of invariant probability measures) and positive Harris recurrence 61
4.2.2 Criterion for finite expectations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
4.2.3 Criterion for recurrence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65
4.2.4 Criterion for transience . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66
4.2.5 Criterion for almost sure convergence to equilibrium . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
4.2.6 State dependent drift criteria: Deterministic and random-time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68
4.2.7 Convergence Rates to Equilibrium . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
4.3 Applications to Stochastic Learning Algorithms and Iterative Dynamics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74
4.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76
4.5 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77

5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming . . . . 81
5.1 Dynamic Programming, Optimality of Markov Policies and Bellman’s Principle of Optimality . . . . . . . . . . . 81
5.1.1 Backwards Induction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82
5.1.2 Optimality of Deterministic Markov Policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
5.1.3 Bellman’s principle of optimality and Dynamic Programming . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
5.1.4 Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
5.2 Existence of Minimizing Selectors and Measurability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
5.2.1 Some Relaxations on the Measurable Selection Conditions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88
5.3 The Linear Quadratic Regulator (LQR) Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
5.4 Optional: A Strategic Measures Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
5.5 Infinite Horizon Optimal Discounted Cost Control Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94
5.5.1 Value Iteration Algorithm and Regularity of Value Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99
5.5.2 Lipschitz Regularity of Value Functions and the Case with Unbounded Costs . . . . . . . . . . . . . . . . . . . 102
5.6 Regularity of Transition Kernels and Optimal Value Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104
5.7 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106

6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter . . . . . . . . . . . . 111
6.1 Enlargement of the State-Space and the Construction of a Controlled Markov Chain . . . . . . . . . . . . . . . . . . . 111
6.2 The Linear Quadratic Gaussian (LQG) Problem and Kalman Filtering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
6.2.1 A Supporting Result on Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
vi Contents

6.2.2 The Linear Quadratic Gaussian Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114


6.2.3 Estimation and Kalman Filtering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115
6.2.4 Optimal Control of Partially Observed LQG Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118
6.3 On the Controlled Markov Construction in the Space of Probability Measures and Extension to General
Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
6.3.1 Non-linear Filter in the Standard Borel setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
6.3.2 Continuity Properties of Belief-MDP: Weak Continuity and Wasserstein Continuity of Filter
Kernels . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124
6.3.3 Existence of Optimal Policies: Discounted Cost and Average Cost . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
6.3.4 A useful structural result: Concavity of the value function in the priors . . . . . . . . . . . . . . . . . . . . . . . . 128
6.4 Filter Stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
6.5 Bibliographic Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
6.6 Appendix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
6.6.1 Proof of Theorems 6.3.4 and 6.3.3. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
6.7 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139

7 The Average Cost Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141


7.1 Average Cost and the Average Cost Optimality Equation (ACOE) or Inequality (ACOI) . . . . . . . . . . . . . . . . . 141
7.2 The Value Iteration and Contraction Approach to the Average Cost Problem . . . . . . . . . . . . . . . . . . . . . . . . . . 145
7.2.1 Contraction under the span semi-norm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145
7.2.2 Contraction under sup norm via minorization by equivalence with a discounted cost problem . . . . . . 147
7.3 The Vanishing Discounted Cost Approach to the Average Cost Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147
7.3.1 Finite state and action spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147
7.3.2 Standard Borel state and action spaces, ACOE and ACOI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148
7.4 The Convex Analytic Approach to Average Cost Markov Decision Problems . . . . . . . . . . . . . . . . . . . . . . . . . . 154
7.4.1 Finite state/action setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155
7.4.2 General state/action spaces under weak continuity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157
7.4.3 General state/action spaces under strong continuity in actions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
7.4.4 Optimality of deterministic stationary policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162
7.4.5 Sample-path optimality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163
7.5 Constrained Markov Decision Processes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166
7.6 Bibliographic Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
7.7 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167

8 Numerical and Approximation Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171


8.1 Value and Policy Iteration Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171
Contents vii

8.1.1 Value Iteration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171


8.1.2 Policy Iteration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171
8.1.3 Receding Horizon Algorithms / Model Predictive Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174
8.2 Approximation through Quantization of the State and the Action Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174
8.2.1 Finite Action Approximation to MDPs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174
8.2.2 Finite State Approximation to MDPs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 177
8.2.3 Finite Model MDP Approximation: Quantization of both the State and Action Spaces . . . . . . . . . . . . 180
8.3 Numerical Methods for POMDPs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
8.3.1 Near optimality of quantized policies under weak Feller or Wasserstein regularity of non-linear
filters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
8.3.2 Near-optimality of finite window policies under filter stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
8.4 Bibliographic Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185
8.5 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 186

9 Reinforcement Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 189


9.1 Stochastic Learning Algorithms and the Q-Learning Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 189
9.1.1 Q-Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 189
9.1.2 Reinforcement Learning for the Average Cost Criterion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
9.1.3 Synchronous Q-Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
9.2 Reinforcement Learning Methods for POMDPs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
9.2.1 Near optimality of quantized policies under weak Feller property of non-linear filters . . . . . . . . . . . . 195
9.2.2 Near-optimality of finite window policies under filter stability and Q-learning convergence . . . . . . . 195
9.3 Q-Learning For Continuous State and Action Spaces: Quantized Q-Learning, its Convergence and
Near-Optimality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199
9.3.1 Error Analysis for Convergence of Quantized Q-Learning for Continuous Space MDPs . . . . . . . . . . 201
9.4 A General Q-Learning Convergence Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 203
9.5 Bibliographic Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 204
9.6 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205

10 Decentralized and Multi-Agent Stochastic Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209


10.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
10.2 Solution Concepts, Information Structures and Witsenhausen’s Intrinsic Model . . . . . . . . . . . . . . . . . . . . . . . . 210
10.2.1 Witsenhausen’s intrinsic model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
10.2.2 Solution concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211
10.2.3 Classification of information structures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 212
10.2.4 A state space model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213
viii Contents

10.3 Solutions to Static Teams . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215


10.4 Static Reduction of Dynamic Teams: Policy-Dependent and Policy-Indepent Reductions . . . . . . . . . . . . . . . . 219
10.4.1 Static reduction I: Dynamic teams with quasi-classical information structures and their
policy-dependent static reduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 219
10.4.2 Static reduction II: Non-classical information structures and their policy-independent reduction . . . 220
10.4.3 Equivalent static reductions preserve optimality but may not person-by-person-optimality or
stationarity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223
10.4.4 All stochastic dynamic teams are nearly static (with independent measurements) reducible . . . . . . . . 224
10.5 Expansion of information Structures: A recipe for identifying sufficient information . . . . . . . . . . . . . . . . . . . . 224
10.6 Convexity of Decentralized Stochastic Control Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 224
10.6.1 Convexity of static team problems and an equivalent representation of cost functions . . . . . . . . . . . . 225
10.6.2 Convexity of Sequential Dynamic Teams . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 227
10.6.3 Symmetric Team Problems: Optimality of Symmetric Policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 228
10.7 The Strategic Measures Approach to Decentralized Stochastic Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 228
10.7.1 Measurable policies as a subset of randomized policies and strategic measures . . . . . . . . . . . . . . . . . . 228
10.7.2 Sets of strategic measures for static teams . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 229
10.7.3 Sets of strategic measures for dynamic teams in the absence of static reduction . . . . . . . . . . . . . . . . . 231
10.7.4 Measurability properties of sets of strategic measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 232
10.8 Existence of Optimal Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 232
10.8.1 Some Applications and Revisiting Existence Results for Classical (Single-DM) Stochastic Control . 233
10.9 Approximation of Optimal Solutions via Finite Approximations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 235
10.10Dynamic Programming and Centralized MDP Reduction Approaches to Team Decision Problems . . . . . . . . 238
10.10.1Dynamic programming approach based on Common Information and a Controlled Markov State . . 238
10.10.2A Universal Dynamic Program . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 240
10.11Bibliographic Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 241
10.12Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 241

11 Controlled Stochastic Differential Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 245


11.1 Continuous-time Markov processes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 246
11.1.1 Two ways to construct a continuous-time Markov process . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 246
11.1.2 The Brownian motion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 246
11.2 Stochastic Integration, the Itô Integral and Stochastic Differential Equations . . . . . . . . . . . . . . . . . . . . . . . . . . 248
11.2.1 Some subtleties on stochastic integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 248
11.2.2 The Itô Integral . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249
11.2.3 The Itô Formula . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 252
Contents ix

11.2.4 Stochastic Differential Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 253


11.2.5 Some Properties of SDEs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 254
11.2.6 Fokker-Planck equation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 256
11.2.7 Rough Integration [Optional] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 256
11.3 Controlled Stochastic Differential Equations and the Hamilton-Jacobi-Bellman Equation . . . . . . . . . . . . . . . 257
11.3.1 Revisiting the deterministic optimal control problem in continuous-time . . . . . . . . . . . . . . . . . . . . . . . 257
11.3.2 The stochastic case and classes of admissible policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 260
11.3.3 Discounted Infinite Horizon Cost Criterion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 263
11.3.4 Average-Cost Infinite Horizon Cost Criterion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 264
11.3.5 Control up to an Exit Time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 265
11.4 Partially Observed Case, Girsanov’s Theorem and Separated Policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 265
11.4.1 Non-linear filtering in continuous time and Zakai’s equation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 266
11.5 Existence of Optimal Policies under Full, Partial and Decentralized Information . . . . . . . . . . . . . . . . . . . . . . . 267
11.5.1 A related existence discussion in deterministic continuous-time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 267
11.5.2 Existence of Optimal Policies for Fully Observed Stochastic Models . . . . . . . . . . . . . . . . . . . . . . . . . . 269
11.5.3 Existence of Optimal Policies for Partially Observed Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 272
11.5.4 Existence for Models with Decentralized Information . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 274
11.6 Near Optimality of Control Policies Designed for Discrete-time Models via Sampling . . . . . . . . . . . . . . . . . . 277
11.6.1 Fully Observed Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
11.6.2 An alternative discrete-time approximation: Euler-Maruyama discretization . . . . . . . . . . . . . . . . . . . . 280
11.6.3 Borkar’s Control Topology and a Partial Differential Equations Approach . . . . . . . . . . . . . . . . . . . . . . 281
11.7 Stochastic Stability of Diffusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 281
11.8 The Wong-Zakai Theorem and Robustness of the Stratonovich Integral . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 284
11.9 Bibliographic Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 285
11.10Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 285

12 Robustness to Incorrect Models and Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 291


12.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 291
12.1.1 Some Examples and Convergence Criteria for Transition Kernels . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293
12.1.2 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 294
12.2 Continuity and Robustness of Optimal Cost in Convergence of Models (POMDP Case) . . . . . . . . . . . . . . . . . 296
12.2.1 Continuity of Optimal Cost in Convergence of Models (POMDP Case) . . . . . . . . . . . . . . . . . . . . . . . . 296
12.2.2 Robustness to Incorrect Models (POMDP Case) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
12.3 Continuity and Robustness in the Fully Observed Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 300
x Contents

12.3.1 Weak Convergence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 300


12.3.2 Setwise Convergence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 304
12.3.3 Total Variation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 305
12.4 The Average Cost Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 305
12.4.1 Approximation by finite horizon cost . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 306
12.4.2 Continuity under the convergence of transition kernels . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 307
12.4.3 Robustness to Incorrect Controlled Transition Kernel Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 307
12.5 Applications to Data-Driven Learning and Finite Model Approximations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 311
12.5.1 Application of Robustness Results to Data-Driven Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 311
12.5.2 Application to Approximations of MDPs and POMDPs with Weakly Continuous Kernels . . . . . . . . . 313
12.6 Bibliographic Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 314
12.7 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315

A Basics of Function Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 317


A.1 Normed Linear (Vector) Spaces and Metric Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 317
A.1.1 Banach Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
A.1.2 Hilbert Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 320
A.1.3 Separability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 321

B On the Convergence of Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323


B.1 Limit Events and Continuity of Probability Measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323
B.2 Borel-Cantelli Lemma . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323
B.3 Convergence of Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
B.3.1 Convergence almost surely (with probability 1) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
B.3.2 Convergence in Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
B.3.3 Convergence in Mean-square . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
B.3.4 Convergence in Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324

C Some Remarks on Measurable Selections . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 327

D On Spaces of Probability Measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 331


D.1 Convergence of Sequences of Probability Measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 331
D.2 Some Measurability Results on Spaces of Probability Measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 333
D.3 A Generalized Dominated Convergence Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 333
D.4 The w-s Topology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 334
D.5 Lusin’s Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 334
Contents xi

E Relaxed Control Topologies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 335


E.1 Young Topology on Control Policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 335
E.2 Borkar (Weak∗ ) Topology on Control Policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 336
E.3 Some Properties of Young and Borkar topologies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 337

References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 339
Introduction
In a typical differential equations or a systems course, one learns about the behaviour of a system described by differential
or difference equations. For such systems, under mild regularity conditions, a given initial condition (in the absence of
disturbances) leads to a unique solution/output. Even when one cannot obtain an explicit analytical solution, it is often
possible to establish stability properties of solutions.
In many engineering or applied mathematics areas, one also has the liberty to affect the flow of the system through a control
term. Control theory is concerned with shaping the input-output behaviour of a system by possibly utilizing feedback from
system outputs under various design criteria and constraints. The way control actions or variables are generated based on
the information available at the controller is called the control policy or control law. In deterministic control theory, under
mild conditions, a given initial state and a given control policy uniquely specifies the realized path. Such a policy may
be designed to stabilize a system, under stabilizability conditions; or control a system, under controllability conditions.
The deterministic theory has had tremendous impact and success in many applications with commonly considered criteria
being system stability (e.g. convergence to a point or a set with respect to initial state conditions, or boundedness of the
output corresponding to any bounded input), reference tracking, robustness to incorrect models and disturbances (which
may appear in the system itself or in measurements available at the controller), and optimal control.
However, in many applications, the deterministic theory is not directly applicable, as there may be disturbances in a given
system. In such systems, disturbances may appear in the dynamics of a system or in the information available to the
controller. Some general differences between the deterministic and stochastic setups are the following:

• The criteria on stability for stochastic systems require a different approach since stabilization to a point, or to a compact
set or formation, is often too much to ask in a stochastic system.
• The solution concepts for stochastic systems can be significantly different from those in a deterministic setup.
• For optimization problems, the optimality criteria need to be probabilistic in general.
• Informational aspects of control and decision making in the presence of uncertainty lead to significant mathematical
complexity but also versatility in applications (including those in decentralized setups where multiple decision makers
are present either in a cooperative or in an adversarial context, as in game theory). Such informational dependence of
stochastic control is perhaps what particularly distinguishes the stochastic theory from its deterministic counterpart.

Nonetheless, we will see that many concepts and principles from deterministic control theory carry over to the stochastic
setup. For a stochastic system, we will see that even though a control policy and an initial condition does not uniquely
determine the path that a controlled process may take, the probability measure on the future paths is uniquely specified
given a policy. Likewise, the concepts of stability, optimality and observability will all find corresponding interpretations
(though with significant generalizations, refinements, but also limitations). Results from geometric control theory and
robust control theory will lead to remarkable insights.
However, these connections require a strong foundation on probability (and several other areas of mathematics and engi-
neering): before we proceed with the technical study of the subject, which will also touch on the aforementioned application
areas, in the first chapter of these notes a concise but sufficiently detailed review of probability theory will be presented.
Some application areas include: optimal regulation and tracking; optimal filtering of noisy measurements with respect to a
hidden dynamical system and control of such systems; operations research; mathematical finance and optimal investment;
stochastic and data-driven learning methods for optimization (including reinforcement and stochastic learning theoretic
problems and applications); stability and optimization of communication networks (e.g. in optimal routing and scheduling);
information theory (in particular for setups involving causality and feedback); robust design of control systems under
approximation errors, incorrect models and priors; stability analysis and stabilization of stochastic dynamical systems;
decentralized stochastic control of systems; stochastic control in the presence of adverse decision makers (as in stochastic
game theory); and stochastic networked control (control under information constraints between various components of a
control system).
In the lecture notes, following a review chapter on probability, we will first proceed with stochastic stability, optimization
under various criteria, the problems with partial information, and stochastic learning theory. A basic course in stochastic
2 Contents

control could cover the topics mentioned so far. If further time is available, the additional material presented on decen-
tralized stochastic control, stochastic control in continuous-time, and robustness to incorrect models and learning, can be
covered.
1

Review of Probability

1.1 Introduction

Before discussing controlled Markov chains, we first discuss some preliminaries about probability theory.
Many events in the physical world are uncertain; that is, with a given prior knowledge (such as an initial condition)
regarding a process, the future values of the process are not exactly predictable. Probability theory attempts to develop an
understanding for such uncertainty in a consistent way given a number of properties to be satisfied.
Examples of stochastic processes include: a) The temperature in a city at noon throughout some October: This process
takes values in R31 , b) The sequence of outputs of a communication channel modeled by an additive scalar Gaussian noise
when the input sequence is given by x = {x1 , . . . , xn } ∈ Rn (the output process lives in Rn ), c) Infinite copies of a
discrete-time coin flip process (living in {H, T }Z+ , where H denotes the head and T denotes the tail outcome), d) The
trajectory of a plane flying from point A to point B (taking values in C([t0 , ∞); R3 ), the space of all continuous paths in
R3 with xt0 = A, xtf = B for some t0 < tf ∈ R), e) The exchange rate between the Canadian dollar and the American
dollar on a given time index T .
Some of these processes take values in countable spaces, some do not. If the state space X in which a random variable
takes values is finite or countably infinite, it suffices to associate with each point x ∈ X a number which determines the
likelihood of the event that x is the realized value of the variable. However, when X is uncountable, only focusing on such
realizations is not sufficiently descriptive and further technical intricacies arise. Accordingly, the notion of an event needs
to be carefully defined. First, if some event A takes place, it must be that the complement of A (that is, this event not
happening) must also be defined. Furthermore, if A and B are two events, then the intersection must also be an event. This
line of thought will motivate us for a more formal analysis below. In particular, one needs to construct probability values by
first defining values for certain events and extending such probabilities to a larger class of events in a consistent fashion (in
particular, one does not first associate probability values to single points as we do in countable state spaces). These issues
are best addressed with a precise characterization of probability and random variables.
Probability theory is a versatile mathematical construction to model and study uncertainty in the real world. In the follow-
ing, we will present a rigorous, though concise, review of probability. For a more complete exposition the reader could
consult with several comprehensive texts on probability theory, such as [43, 57, 110, 129, 316] and texts on stochastic
processes, such as [150, 157, 332].

1.2 Measures and Integration

Let X be a collection of points. Let F be a collection of subsets of X with the following properties such that F is a σ-field
(also called a σ-algebra), that is:

• X∈F
4 1 Review of Probability

• If A ∈ F, then X \ A ∈ F
S∞
• If Ak ∈ F, k = 1, 2, 3, . . . , then k=1 Ak ∈ F (that is, the collection is closed under countably many unions).

By De Morgan’s laws, and set properties, it can be shown that the collection has to be closed under countable intersections
as well.
For example, the full power-set of any set is a σ-field.
If the third item above holds for only finitely many unions or intersections, then the collection of subsets is said to be a
field or algebra over X.
With the above, (X, F) is termed a measurable space (that is we can associate a measure to this space; which we will
discuss shortly).

Remark 1.1. Subsets in σ-fields can be interpreted to represent information that a controller has with regard to an underlying
process. We will discuss this interpretation further and this will be a recurring theme in our discussions in the context of
stochastic control.

A σ-field J is generated by a collection of sets A, if J is the smallest σ-field containing the sets in A, and in this case, we
write J = σ(A).

Exercise 1.2.1 Let X = {a, b, c}. (i) Find σ({a}). (ii) Find σ({a}, {b}, {c}).

We consider an important special case in the following.

1.2.1 Borel σ-field

An important class of σ-fields is the Borel σ-field on a metric (or more generally, topological) space. Such a σ-field is
the one which is generated by open sets. The term open naturally depends on the space being considered. For this course,
we will mainly consider spaces which are complete, separable and metric spaces (such as the space of real numbers R, or
countable sets) see Appendix A). Recall that in a metric space with metric d, a set U is open if for every x ∈ U , there exists
some ϵ > 0 such that {y : d(x, y) < ϵ} ⊂ U . We note also that the empty set is a special open set.
The Borel σ-field on R is then the one generated by sets of the form (a, b) ⊂ R, that is, open intervals (it is important to
note here that every open set in R can be expressed a union of countably many open intervals). It is also important to note
that not all subsets of R are Borel sets, that is, elements of the Borel σ-field; see e.g. Exercise 1.6.7.
We will denote the Borel σ-field on a space X as B(X).

Exercise 1.2.2 Show that for every a ∈ R, {a} ∈ B(R), by writing a = ∩n∈N (a − n1 , a + n1 ).

We can also define a Borel σ-field on a product space. Let X be a complete, separable, metric space (with metric d). Let
XZ+ denote the infinite product of X so that x = (x0 , x1 , x2 , · · · ) ∈ XZ+ , where xk ∈ X for k ∈ Z+ . If this space is
P∞ d(xi ,yi )
endowed with the product metric (such a metric is defined as: ρ(x, y) = i=0 2−i 1+d(x i ,yi )
, x, y ∈ XZ+ ), sets of the
Q
form i∈Z+ Ai , where only finitely many of these sets are not equal to X and these sets are open; and unions of such sets
form open sets. We define cylinder sets in this product space as:

B[Am ,m∈I] = {x ∈ XZ+ , xm ∈ Am , m ∈ I},

with Am ∈ B(X) and where I ⊂ Z with |I| < ∞, that is, the set I has finitely many elements. Thus, in the above,
if x ∈ B[Am ,m∈I] , then, xm ∈ Am for m ∈ I and the remaining terms (that is, the xm values for m ∈ / I) can be
taken arbitrarily from X. We can thus view a cylinder set as a pre-image of the projection operation onto finitely many
coordinates. The σ-field generated by such open cylinder sets is the Borel σ-field on the product space. Such a construction
1.2 Measures and Integration 5

is important for stochastic processes (and is the reason why while studying certain properties of stochastic processes one
often only considers finite dimensional distributions).

Remark 1.2. A space which admits a metric under which it is complete and separable is called a Polish space; while the
metric is often not apriori specified for such a space, we will often assume that a metric is given and define a Polish metric
space to be a complete separable and metric space. A Borel subset of a Polish space is called a standard Borel space [293].
A very important fact is that any Polish space is related to either a finite set, or a countably infinite set, or R, through a
bijection (that is, via a measurable function -to be defined shortly- with a measurable inverse).

1.2.2 Measurable Function

If (X, F) and (Y, G) are measurable spaces; we say a mapping from h : (X, F) → (Y, G) is a measurable function if

h−1 (B) := {x : h(x) ∈ B} ∈ F ∀B ∈ G

In the particular case involving Borel σ-fields, if (X, B(X)) and (Y, B(Y)) are measurable spaces; we say a mapping from
h : X → Y is (Borel) measurable if

h−1 (B) = {x : h(x) ∈ B} ∈ B(X), ∀B ∈ B(Y)

Theorem 1.2.1 To show that a function is measurable, it is sufficient to check the measurability of the inverses of sets that
generate the σ-algebra on the image space.

See Section 1.5.1 for a proof. Therefore, for Borel measurability, it suffices to check the measurability of the inverse images
of open sets. Furthermore, for real valued functions, to check the measurability of the inverse images of open sets, it suffices
to check the measurability of the inverse images sets of the form {(−∞, a], a ∈ R}, {(−∞, a), a ∈ R}, {(a, ∞), a ∈ R}
or {[a, ∞), a ∈ R}, since each of these generate the Borel σ-field on R. In fact, here we can restrict a to be Q-valued,
where Q is the set of rational numbers (since {x : x < r} = ∪q∈Q,q<r {x : x < q}; often this reasoning is why we call
such sigma-fields countably generated).
It is instructive to view measurability in terms of informativeness of a σ-field. Let X = {a, b, c} and let F1 be as in
Exercise 1.2.1(i) and F2 be as in Exercise 1.2.1(ii). Let Y = {0, 1} and G = σ({0}, {1}). Now, let F : X → Y be a
map so that F −1 (0) = {a, b} and F −1 (1) = {c}. Then, we can conclude that this map defines a measurable function
from (X, F2 ) → (Y, G) but it is not a measurable function from (X, F1 ) → (Y, G): the reason is that the information on
whether F (x) = 1 (that is x = c) is not an element of F1 (and thus, this information that x = c or not, is not available as
information under F1 ).

1.2.3 Measure

Let (X, F) be a measurable space. A positive measure µ on (X, F) is a map from F to [0, ∞] which is countably additive
such that for Ak ∈ F and Ak ∩ Aj = ∅:
  X ∞
µ ∪∞ k=1 A k = µ(Ak ).
k=1

Definition 1.2.1 µ is a probability measure if it is positive and µ(X) = 1.

Definition 1.2.2 A measure µ is finite if µ(X) < ∞, and σ-finite if there exist a collection of subsets Ak ∈ F such that
X = ∪∞ k=1 Ak with µ(Ak ) < ∞ for all k.
6 1 Review of Probability

On the real line R, the Lebesgue measure λ is defined on the Borel σ-field (in fact on a somewhat larger field obtained
through adding all subsets of Borel sets of measure zero: this is known as completion of a σ-field) such that for A = (a, b),
λ(A) = b−a. Borel σ-field of subsets is a strict subset of Lebesgue measurable sets, that is there exist Lebesgue measurable
sets which are not Borel sets. For a definition and the construction of Lebesgue measurable sets, see [43]. Countable
subsets of R all have zero Lebesgue measure but there also exist Lebesgue measurable sets of measure zero which contain
uncountably many elements, for a well-studied example see the Cantor set [110]. Not all subsets of R are Lebesgue
measurable (and thus, not Borel either), see e.g. Exercise 1.6.7.

1.2.4 The Extension Theorem (Optional)

Theorem 1.2.2 [The Extension Theorem (Carathéodory)] Let M be an algebra over X, and suppose that there exists a
map (called a pre-measure) P : M → R+ so that for any (possibly P countably infinitely many) pairwise disjoint sets
An ∈ M, if the countable union ∪n An ∈ M, then P (∪n An ) = n P (An ). Suppose also that there exists a countable
collection of sets Bn with X = ∪n Bn , each with P (Bn ) < ∞ (that is P is σ-finite). Then, there exists a unique measure
P ′ on the σ-field generated by M, σ(M), which is consistent with P on M.

The above is useful since, when one states that two measures are equal it suffices to check whether they are equal on the
algebra of sets which generate the σ-field, and not necessarily on the entire σ-field. More importantly, a refinement of the
above can be used to define or construct a measure on a σ-field, such as the Lebesgue measure on B(R).
The following is a refinement useful for stochastic processes. It, in particular, does not require a pre-measure defined apriori
before an extension [4]:

Theorem 1.2.3 [Kolmogorov’s Extension Theorem] Let X be a complete and separable metric space, and for all n ∈ N
let µn be a sequence of probability measures on Xn , the n product of X, such that

µn (A1 × A2 × · · · × An ) = µn+1 (A1 × A2 × · · · × An × X),

every sequence of Borel sets Ak ⊂ X. Then, there exists a unique probability measure µ on (XN , B(XN )) which is consistent
with each of the µn ’s.

A further related result, which often in stochastic control is cited in the context of extensions, is the Ionescu-Tulcea Ex-
tension Theorem [164, Appendix C]; where conditional probability measures (stochastic kernels) are defined (instead of
probability measures on finite dimensional product spaces) as a starting assumption, before an extension to the infinite
product space is established.
Thus, if the σ-field on a product space is generated by the collection of finite dimensional cylinder sets, one can define a
measure in the product space which is consistent with the finite dimensional distributions.
Likewise, we can construct the Lebesgue measure on B(R) by defining it on finitely many unions and intersections of
intervals of the form (a, b), [a, b), (a, b] and [a, b], and the empty set, thus forming an algebra (or a field), and extending
this to the Borel σ-field. Thus, the relation µ(a, b) = b − a for b > a is sufficient to define the Lebesgue measure.

Remark 1.3. A related general result is as follows: Let S be a σ-field. A class of subsets A ⊂ S is called a separating class
if two probability measures that agree on A agree on the entire S. A class of subsets is a π-system if it is closed under finite
intersections. The class A is a separating class if it is both a π-system and it generates the σ-field S; see [42] or [43].

1.2.5 Integration

Let h be a non-negative measurable function from (X, B(X)) to (R, B(R)). The Lebesgue integral of h with respect to a
measure µ can be defined in three steps:
1.2 Measures and Integration 7

First, for A ∈ B(X), define 1{x∈A} (or 1(x∈A) , or 1A (x)) as an indicator function for event x ∈ A, that is the value that
the function takes is 1 if x ∈ A, and 0 otherwise. In this case, define
Z
1{x∈A} µ(dx) := µ(A).
X

Pn functions h such that, there exist A1 , A2 , . . . , An all in B(X) and positive numbers b1 , b2 , . . . , bn
Now, let us define simple
such that hn (x) = k=1 bk 1{x∈Ak } . For such functions, define
Z n
X
hn (x)µ(dx) := bk µ(Ak ).
X k=1

Now, for any given measurable h, there exists a sequence of simple functions hn such that hn (x) ↑ h(x) monotonically,
that is hn+1 (x) ≥ hn (x) (for a construction, if h only takes non-negative values, consider partitioning the positive real line
to two intervals [0, n) and [n, ∞), and partition [0, n) to n2n uniform intervals, define hn (x) to be the lower floor of the
interval that contains h(x): thus

hn (x) = k2−n , if k2−n ≤ h(x) < (k + 1)2−n , k = 0, 1, · · · , n2n − 1,

and hn (x) = n for h(x) ≥ n. By definition, and since h−1 ([k2−n , (k + 1)2−n )) is Borel, hn is a simple function. If
the function takes also negative values, write h(x) = h+ (x) − h− (x), where h+ is the non-negative part and −h− is the
negative part, and construct the same for h− (x). We define the limit (which exists as a real valued monotonically increasing
sequence) as the Lebesgue integral: Z Z
lim hn (x)µ(dx) =: h(x)µ(dx)
n→∞
R R R
We note that the notation hdµ or h(x)dµ(x) can also be used in place of h(x)µ(dx).

1.2.6 Fatou’s Lemma, the Monotone Convergence Theorem and the Dominated Convergence Theorem

Theorem 1.2.4 (Monotone Convergence Theorem) If µ is a σ-finite positive measure on (X, B(X)) and {fn , n ∈ Z+ }
is a sequence of measurable functions from X to R which pointwise, monotonically, converges to f so that 0 ≤ fn (x) ≤
fn+1 (x) for all n, and
lim fn (x) = f (x),
n→∞

for µ−almost every x, then Z Z


f (x)µ(dx) = lim fn (x)µ(dx)
X n→∞ X

The following is a consequence of the monotone convergence theorem, but is a critical result which will be utilized later in
the upcoming chapters.

Theorem 1.2.5 (Fatou’s Lemma) If µ is a σ-finite positive measure on (X, B(X)) and {fn , n ∈ Z+ } is a sequence of
measurable functions, bounded from below, from X to R, then
Z Z
lim inf fn (x)µ(dx) ≤ lim inf fn (x)µ(dx)
X n→∞ n→∞ X

Theorem 1.2.6 (Dominated Convergence Theorem) If (i) µ is a σ-finite positive measure on (X, B(X)), (ii) g is a Borel
measurable function with Z
g(x)µ(dx) < ∞,
X
8 1 Review of Probability

and (iii) {fn , n ∈ Z+ } is a sequence of measurable functions from X to R which satisfy |fn (x)| ≤ g(x) for µ−almost
every x, and limn→∞ fn (x) = f (x), then:
Z Z
f (x)µ(dx) = lim fn (x)µ(dx)
X n→∞ X

Note that for the monotone convergence theorem, there is no restriction on boundedness; whereas for the dominated
convergence theorem, there is a boundedness condition. On the other hand, for the dominated convergence theorem, the
pointwise convergence does not have to be monotone.
There also exist generalized versions of these theorems, where the measures themselves are time-varying, but converge
to a limit measure is in some appropriate sense; see in particular Theorem D.3.1 (building on [211, 287]). These will be
discussed later in further detail (and will be seen to be particularly important for stochastic control applications, and in
particular on robust stochastic control and learning theory).

1.3 Probability Space and Random Variables

Let (Ω, F) be a measurable space. If P is a probability measure, then the triple (Ω, F, P ) is called a probability space.
Here Ω is a set called the sample space. F is called the event space, and this is a σ-field of subsets of Ω.
Let (E, E) be another measurable space and X : (Ω, F, P ) → (E, E) be a measurable map. We call X an E−valued
random variable. The image under X defines a probability measure on (E, E), called the law of X.
The σ-field generated by the events {{w : X(w) ∈ A}, A ∈ E}, that is {X −1 (A), A ∈ E}, is called the σ-field generated
by X and is denoted by σ(X).
Consider a coin flip process, with possible outcomes {H, T }, heads or tails. We have a good intuitive understanding on
the environment when someone tells us that a coin flip leads to the value H with probability 21 . Based on the definition of
a random variable, we view then a coin flip outcome as a deterministic function from some space (Ω, F, P ) to the binary
output space consisting of a head and a tail event. Here, P denotes the uncertainty measure (you may think of the initial
condition of the coin when it is being flipped, the flow dynamics in the air, the conditions on the surface where the coin
touches etc.; we encode all these aspects and all the uncertainty in the universe with the abstract space (Ω, F, P )). You can
view then the σ-field generated by such a coin flip as a partition of Ω: if certain things take place the outcome is a H and
otherwise it is a T and the outcomes give us information on (the state of) the universe.
A useful fact about measurable functions (and thus random variables) is the following result.

Theorem 1.3.1 Let fn be a sequence of measurable functions from (Ω, F) to a complete separable metric space (X, B(X).
Then, lim supn→∞ fn (x), lim inf n→∞ fn (x) are measurable. In particular, if f (x) = limn→∞ fn (x) exists, then f is
measurable.

Similar to Theorem 1.2.1, this theorem implies that to verify whether a real valued mapping f is a Borel measurable
function, it suffices to check if f −1 (a, b) ∈ B(R) for a < b since one can construct a sequence of simple functions which
will converge to any measurable f , as discussed earlier. It suffices then to check if f −1 (−∞, a) ∈ B(R) for a ∈ R.

1.3.1 More on Random Variables and Probability Density Functions

Consider a probability space (X, B(X), P ) and consider an R-valued random variable U measurable with respect to
(X, B(X)).
This random variable induces a probability measure µ on B(R) such that for some (a, b] ∈ B(R):
 
µ((a, b]) = P (U ∈ (a, b]) = P {x : U (x) ∈ (a, b]} = P (U −1 ((a, b])
1.3 Probability Space and Random Variables 9

When U is R−valued, the expectation of U is given with


Z
E[U ] = µ(dx)x,
R

whenever this is defined (i.e., E[|U |] < ∞, in which


R xcase we say that U is integrable). We define F (x) = µ(−∞, x] as the
cumulative distribution function of U . If F (x) = −∞ p(s)λ(ds) for some p, p is called the probability density function
(with respect to the Lebesgue measure) of µ. If such a density function exists, we can then write
Z
E[U ] = p(x)xdx
R

If a probability density function p exists, the measure µ is said to be absolutely continuous with respect to the Lebesgue
R function p is the Radon-Nikodym derivative of µ with respect to the Lebesgue measure λ
measure. In particular, the density
in the sense that for all Borel A: A p(x)λ(dx) = µ(A). A probability density function does not always exist. In particular,
whenever there is a probability mass on a given point, then a probability density function does not exist; hence in R, if for
some x, µ({x}) > 0, then we say there is a probability mass at x, and a density function does not exist.
However, one can also consider density functions with respect to more general positive measures (that is, different from
the Lebesgue measure), we will consider such conditions later in the notes. If X is countable, we can write P ({x = m}) =
p(m), where p is called the probability mass function; this can be viewed as a density with respect to the (discrete) counting
measure.
Some examples of commonly encountered random variables, with their probability density or mass functions, are as fol-
lows:

• Gaussian (with mean µ and variance σ 2 : N (µ, σ 2 )):

1 −(x−µ)2
p(x) = √ e 2σ2 x∈R
2πσ

• Exponential (with parameter λ):

F (x) = 1 − e−λx , p(x) = λe−λx x ∈ R+

• Uniform on [a, b] (U ([a, b])):


x−a 1
F (x) = , p(x) = x ∈ [a, b]
b−a b−a

• Poisson with rate λ > 0 on Z+

λm e−λ
p(m) = , m ∈ Z+
m!

• Binomial (B(n, p)):  


n k
p(k) = p (1 − p)n−k k ∈ {0, 1, · · · , n}
k

If n = 1, we also call a binomial variable a Bernoulli variable.

1.3.2 Independence and Conditional Probability

Consider A, B ∈ B(X) such that P (B) > 0. The quantity


10 1 Review of Probability

P (A ∩ B)
P (A|B) =
P (B)

is called the conditional probability of event A given B. The measure P (·|B) defined on B(X) is itself a probability
measure. If
P (A|B) = P (A),
A and B are said to be independent events. A countable collection of events {An } is independent if for any finitely many
sub-collections Ai1 , Ai2 , . . . , Aim , we have that

P (Ai1 , Ai2 , . . . , Aim ) = P (Ai1 )P (Ai2 ) . . . P (Aim ).

Here, we use the notation P (A, B) = P (A ∩ B). A sequence of events is said to be pairwise independent if for any two
pairs (Am , An ): P (Am , An ) = P (Am )P (An ). Pairwise independence is a weaker concept than independence, that is
there exist examples where a collection of random variables is pairwise independent but not independent.
Conditional probability and expectation will be discussed in more detail later in Chapter 4.

1.4 Stochastic Processes and Markov Chains

One can define a sequence of random variables as a single random variable living in a product space; that is, we can
consider {x1 , x2 , · · · , xN , · · · } as an individual random variable X which is an XZ+ -valued random variable, where now
the events are to be defined on the product space.
Let X be a complete, separable, metric space and let T = Z or T = Z+ . Let B(X) denote the Borel sigma-field over X. Let
Σ = XT denote the sequence space of all one-sided (with T = Z+ ) or two-sided (with T = Z) infinitely many random
variables drawn from X. Thus, if T = Z, x ∈ Σ then x = {. . . , x−1 , x0 , x1 , . . . } with xi ∈ X, i ∈ T . Let Xn : Σ → X
denote the coordinate function such that Xn (x) = xn . Let B(Σ) denote the smallest sigma-field containing all cylinder
sets of the form {x : Xi (x) = xi ∈ Bi , m ≤ i ≤ n} where Bi ∈ B(X), for all integers m, n. We can define a probability
measure by a characterization on these finite dimensional cylinder sets, by (the extension) Theorem 1.2.3.
A similar characterization also applies for continuous-time stochastic processes, where T is uncountable. The extension
requires more delicate arguments, since finite-dimensional characterizations are too weak to uniquely define a sigma-field
on a space of continuous-time paths which is consistent with such distributions. Such technicalities arise in the discussion
for continuous-time Markov chains and controlled processes, typically requiring a construction where realizations take
values from a separable product space (such as the space of continuous sample paths; in this case, the sample path values
on a countably dense subset uniquely determine the entire sample path and hence the discrete-time theory, essentially, is
applicable); see Section 11.1.1 for further discussion.
In much of these notes, our focus will primarily be on discrete-time processes; however, we will note later that the analysis
for continuous-time processes essentially follows from similar constructions with further structures that one needs to im-
pose on continuous-time processes (such as some continuity properties of the sample paths). Further discussion on this is
presented in Chapter 11.

1.4.1 Markov Chains

If the probability measure on an XZ+ -valued sequence is such that for every k ∈ N, for every Borel Ak+1 and (P -almost
surely) all realizations x[0,k] ,

P (xk+1 ∈ Ak+1 |xk , xk−1 , . . . , x0 ) = Pk (xk+1 ∈ Ak+1 |xk ),

for some conditional probability measure Pk , then {xk } is said to be a Markov chain. If Pk is a constant and does not
depend on k, the chain is said to be a time-homogeneous chain, otherwise it is time-inhomogeneous. Thus, for a Markov
chain, the immediate state is sufficient to predict the future (and past variables are not needed).
1.6 Exercises 11

One way to construct a Markov chain is via the following: Let {xt , t ≥ 0} be a random sequence with state space
(X, B(X)), and defined on a probability space (Ω, F, P), where B(X) denotes the Borel σ-field on X, Ω is the sam-
ple space, F a sigma field of subsets of Ω, and P a probability measure. For x ∈ X and D ∈ B(X), we let
P (x, D) := P(xt+1 ∈ D|xt = x) denote the transition probability from x to D, that is the probability of the event
{xt+1 ∈ D} given that xt = x. Thus, the Markov chain is completely determined by the transition probability and the
probability of the initial state, P (x0 ∈ ·). The
R probability of the event {xt+1 ∈ D} for any t can be computed recursively
by starting at t = 0, with P(x1 ∈ D) = P (x1 ∈ D|x0 = x)P(x0 ∈ dx), and iterating with a similar formula for
t = 1, 2, . . . (building on the Ionescu-Tulcea Extension Theorem [164, Appendix C], which was discussed earlier).
Hence, if the probability of the same event given some history of the past and the present does not depend on the past, and
hence is given by the same quantity regardless of the past realizations as long as the present realization is fixed (almost
surely), the chain is a Markov chain. As an example, consider the following linear system:

xt+1 = axt + wt ,

where {wt } is an independent sequence of random variables for some a ∈ R. The process {xt } is Markov. We also
note that every time-homogeneous Markov chain admits a stochastic, functional and sample-path, realization of the form
xk+1 = f (xk , wk ) where f is measurable and wk is an i.i.d. [0, 1]-valued process (see [143, Lemma 1.2], [56, Lemma
3.1], or [21, Lemma F]). This realization result will be useful later on.
We will continue our discussion on Markov chains after discussing controlled Markov chains in the following chapter.

1.5 Appendix

1.5.1 Proof of Theorem 1.2.1

Observe that set operations satisfy that for any B ∈ B(Y): h−1 (Y \ B) = X \ h−1 (B) and

h−1 (∪∞ ∞
i=1 Bi ) = ∪i=1 h
−1
(Bi ), h−1 (∩∞ ∞
i=1 Bi ) = ∩i=1 h
−1
(Bi ).

Define the set of all subsets of Y whose inverses are Borel

M := {B ⊂ Y : h−1 (B) ∈ B(X)}.

Note that Y ⊂ M and by the discussion above, this set is closed under countably many unions. Thus, this M is a σ-algebra
over Y. Note also that this set contains open sets in Y, by the fact that h is measurable, and since this set contains open sets
(and that B(Y) is the smallest σ-algebra containing open sets), it must be that B(Y) ⊂ M. ⋄

1.6 Exercises

Exercise 1.6.1 a) Let H be some set and for all β ∈ H, Fβ be a σ-field of subsets over some set X. Let
\
F= Fβ
β∈H

Show that F is also a σ-field on X.


For a space X, on which a metric is defined, the Borel σ-field is generated by the collection of open sets. This means that,
the Borel σ-field is the smallest σ-field containing open sets, and as such it is the intersection of all σ-fields containing
open sets.
b) Show that any open set in R under the usual distance d(x, y) = |x − y|, can be written as a countable union of intervals.
A consequence of this result is that, on R, the Borel σ-field is the smallest σ-field containing open intervals.
12 1 Review of Probability

c) Is the set of rational numbers an element of the Borel σ-field on R? Is the set of irrational numbers an element?
d) Let X be a countable set. On this set, let us define a metric as follows:
(
0, if x = y
d(x, y) =
1, if x ̸= y

Show that, the Borel σ-field on X is generated by the collection of singletons {{x}, x ∈ X}; this is the power set, that is,
the set of all subsets of X.
e) Let X = R and consider the metric d defined as above in part d). Is the σ-field generated by open sets according to
this metric the same as the Borel σ-field on R (under the usual distance metric on R)? Finally, consider the σ-algebra
generated by individual points (singletons, that is {{x}, x ∈ X}); is this the same as the Borel σ-field or is this the same
as the power set on R?

Exercise 1.6.2 A Borel subset of a complete, separable and metric (i.e., a Polish) space is called a standard Borel space.
If (X, B(X)) and (Y, B(Y)) are standard Borel spaces; we say a mapping from h : X → Y is (Borel) measurable if

h−1 (B) = {x : h(x) ∈ B} ∈ B(X), ∀B ∈ B(Y)

Prove the following statement: To show that a function h is Borel measurable, it is sufficient to check the measurability of
the inverses (under h) of open sets in Y.

Exercise 1.6.3 Investigate the following limits in view of the convergence theorems.
R1 R1
a) Check if limn→∞ 0 xn dx = 0 limn→∞ xn dx.
R1 R1
b) Check if limn→∞ 0 nxn dx = 0 limn→∞ nxn dx.
R R
c) Define fn (x) = n1{0≤x≤ n1 } . Find limn→∞ fn (x)dx and limn→∞ fn (x)dx. Are these equal?

Exercise 1.6.4 a) Let X and Y be real-valued random variables defined on a given probability space. Show that X 2 and
X + Y are also random variables.
b) Let F be a σ-field of subsets over a set X and let A ∈ F. Prove that {A ∩ B, B ∈ F } is a σ-field over A (that is a
σ-field of subsets of A).
Hint for part a: The following equivalence holds: {X + Y < x} ≡ ∪r∈Q {X < r, Y < x − r}. To check if X + Y is a
random variable, it suffices to check if the event {X + Y < x} = {ω : X(ω) + Y (ω) < x} is an element of F for every
x ∈ R.

Exercise 1.6.5 Let fn be a sequence of measurable functions from (Ω, F) to (R, B(R)). Show f (ω) = lim supn→∞ fn (ω)
and g(ω) = lim inf n→∞ fn (ω) define measurable functions.

Exercise 1.6.6 Let X and Y be real-valued random variables defined on a given probability space (Ω, F, P ). Suppose
that X is measurable on σ(Y ). Show that there exists a function f such that X = f (Y ).
This result also holds if X and Y are standard Borel valued random variables.

Exercise 1.6.7 Consider the interval [0, 1]. We have seen that the Lebesgue measure λ satisfies λ([a, b)) = U ([a, b]) =
b − a for 0 ≤ a ≤ b ≤ 1. Consider now the following question: does every subset S ⊂ [0, 1] admit a Lebesgue measure?
In the following we will provide a counterexample, known as the Vitali set.
Let us define an equivalence class among points in [0, 1] such that x ∼ y if x−y ∈ Q. This equivalence definition partitions
[0, 1] into disjoint sets. Note that there are countably many points in each equivalent class.
1.6 Exercises 13

Let A be a subset which picks exactly one element from each equivalent class (here, we adopt what is known as the Axiom
of Choice [43]). Since A contains an element from each equivalence class, each point of [0, 1] is contained in the union
∪q∈Q (A + q). Furthermore, since A contains only one point from each equivalence class, the sets A + q, for different q,
are disjoint, for otherwise there would be two sets which could include a common point: A + q and A + q ′ would include a
common point, leading to the result that the difference x − q = z and x − q ′ = z are both in A, a contradiction, since there
should be at most one point which is in the same equivalence class as x − q = z. The Lebesgue measure is shift-invariant,
therefore λ(A) = λ(A + q). Observe that [0, 1] ⊂ ∪q∈Q∪[−1,1] {A + q} ⊂ [−1, 2]. Since a countable sum of identical
non-negative elements can either become ∞ or 0, the contradiction follows: We can’t associate a number to this set and
as a result, this set is not a Lebesgue measurable set (and also not a Borel set).
2

Controlled Markov Chains

In the following, we discuss controlled Markov models under a variety of informational and dynamical setups.

2.1 Controlled Markov Models

Consider the following model.

xt+1 = f (xt , ut , wt ), (2.1)

where xt is an X-valued state variable, ut a U-valued control action variable, wt a W-valued an i.i.d noise process, and f
a measurable function. We assume that X, U, W are Borel subsets of complete, separable, metric spaces (such complete,
separable and metric spaces are called Polish metric spaces); such subsets of these spaces are also called standard Borel.
We assume that all random variables live in some probability space (Ω, F, P ).
Using stochastic realization results (see [21, Lemma F], [143, Lemma 1.2], or [56, Lemma 3.1]), it follows that the model
above in (2.1) contains the class of all (X × U)Z+ -valued stochastic processes which satisfy the following probabilistic
characterization: for all Borel sets B ∈ B(X), t ≥ 0, and P -almost all realizations x[0,t] , u[0,t] :

P (xt+1 ∈ B|x[0,t] = a[0,t] , u[0,t] = b[0,t] ) = P (xt+1 ∈ B|xt = at , ut = bt ) =: T (B|at , bt ) (2.2)

where T (·|x, u) is a stochastic kernel from X × U to X (so that for every B, T (B|·, ·) is a measurable function on X × U,
and for every fixed (a, b) ∈ X × U, T (·|a, b) is a probability measure on (X, B(X)). That is, all stochastic processes that
satisfy (2.2) admit a realization in the form (2.1), almost surely. Furthermore, we may take W = [0, 1] without any loss.
Since a system of the form (2.1) satisfies (2.2), it follows that the representations in these equations are equivalent.
A stochastic process which satisfies (2.2) is called a controlled Markov chain.
For the process {xt , ut } to define a stochastic process, in addition to a transition kernel and an initial measure on x0 ,
we need to specify the dependence of ut on the history of the process. Once this is established, through the extension
theorems discussed earlier (and in particular the Ionescu-Tulcea Extension Theorem), one can construct a stochastic process
{xt , ut , t ≥ 0}. This dependence is defined by a control policy.
We start with Fully Observed Controlled Markov Models, otherwise known as Markov Decision Processes (or MDPs).

2.2 Fully Observed Markov Control Problem Model (MDP Models)

A Fully Observed Markov Control Problem, otherwise known as a Markov Decision Process (or MDP), is a five tuple

(X, U, {U(x), x ∈ X}, T , c),


16 2 Controlled Markov Chains

where

• X is the state space, assumed a standard Borel space.


• U is the action space, assumed a standard Borel space.
• K = {(x, u) : u ∈ U(x) ∈ B(U), x ∈ X} is the set of state, control pairs that are feasible. There might be different
states where different control actions are possible/feasible. We will assume that K is standard Borel.
• T is a state transition kernel, that is T (A|x, u) = P (xt+1 ∈ A|xt = x, ut = u), as defined above.
• c : K → R+ is a cost function.

2.2.1 Classes of Control Policies

Admissible Control Policies ΓA

Let H0 := X, Ht = Ht−1 × K for t = 1, 2, . . .. We let It denote an element of Ht , where It = {x[0,t] , u[0,t−1] }. A


deterministic admissible control policy γ is a sequence of functions {γt , t ∈ Z+ } such that γ : Ht → U with ut = γt (It ).
We can also state this as follows: Let us write Ut to emphasize that ut is a realization of the action random variable Ut
under an admissible policy, and likewise let us emphasize that Ht is a random variable with realization It (In the notes, we
will follow this approach of using capital letters when the distinction of whether we are discussing a random variable or its
realization needs to be particularly emphasized explicitly). We say that γt is a function measurable on σ(Ht ) in the sense
that for every Borel B ⊂ U, we have that

{ω : Ut (ω) ∈ B} = Ut−1 (B) ⊂ σ(Ht ).

A randomized admissible control policy is a sequence γ = {γt , t ≥ 0} such that γ : Ht → P(U), with P(U) being the
set of probability measures on U, so that for every realization It , γt (It ) is a probability measure on U. Once again, by
stochastic realization arguments, this is equivalent to writing ut = γt (It , rt ) for some [0, 1]-valued i.i.d. random variable
rt .

Markov Control Policies ΓM

A deterministic Markov control policy γ is a sequence of functions {γt , t ∈ Z+ } with γt : X × Z+ → U such that

ut = γt (xt ),

for each t ∈ Z+ . Hence, the control action only depends on the state and the time, and not the past history. A policy is
randomized Markov if the induced strategic measure satisfies

P γ (ut ∈ C|It ) = γt (ut ∈ C|xt ), C ∈ B(U),

for all t and P γ -almost all xt . Alternatively, we can write ut = γt (xt , rt ) for some [0, 1]-valued i.i.d. random variable rt
and measurable function γt for all t ∈ Z+ .

Stationary Control Policies ΓS

A deterministic stationary control policy γ is a sequence of identical functions {γt , t ∈ Z+ } where for some f : X → U,
γt = f so that
ut = f (xt ).
2.3 Performance Criteria: Optimality and Stability 17

for each t ∈ Z+ .
A policy is randomized stationary if

P γ (ut ∈ C|It ) = f (ut ∈ C|xt ), C ∈ B(U),

for some stochastic kernel f . As earlier, alternatively, we can write ut = f (xt , rt ) for some [0, 1]-valued i.i.d. random
variable rt and measurable function f for all t ∈ Z+ . Hence, the control selection is independent of the past history or
time, given the current state xt .
Often, we will simply identify the stage-wise constant map with the stationary policy γ by an abuse of notation, that is
γ := {γ, γ, · · · }.
As reviewed above and in Chapter 1, according to the Ionescu-Tulcea theorem [164] (or Kolmogorov’s extension theorem),
an initial probability measure µ on X, a transition kernel T , and a control policy γ define a unique probability measure
Pµγ on (X × U)Z+ , which is called a strategic measure [283]. If the initial measure µ is known, sometimes we omit this
subscript while discussing the strategic measure.
PT −1
Consider for now that the objective to be minimized is given by: JN (ν0 , γ) := Eνγ0 [( t=0 c(xt , ut )) + cN (xN )], where
ν0 is the initial probability measure, that is x0 ∼ ν0 . The goal is to find a policy γ ∗ so that

JN (ν0 , γ ∗ ) ≤ JN (ν0 , γ) ∀γ ∈ ΓA .

Such a γ ∗ is called an optimal policy. Here γ can also be called a strategy, or a law.

2.3 Performance Criteria: Optimality and Stability

2.3.1 Several Optimality Criteria and Performance of Policy Classes

Consider a Markov control problem with an objective given as the minimization of


−1
 NX  
JN (ν0 , γ) = Eνγ0 c(xt , ut ) + cN (xN )
t=0

where ν0 denotes the distribution on x0 and cN is a terminal state cost function. For the case with x0 = x, so that ν0 = δx ,
we often simply write
N
X −1 N
X −1
JN (δx , γ) =: JN (x, γ) = Eδγx [ c(xt , ut ) + cN (xN )] = E [γ
c(xt , ut ) + cN (xN )|x0 = x]
t=0 t=0

Such a cost problem is known as an expected finite horizon cost criterion.


We will also consider costs of the following form:

X∞
Jβ (ν0 , γ) = Eνγ0 [ β t c(xt , ut )],
t=0

for some β ∈ (0, 1). This is called an expected discounted infinite horizon cost criterion.
Finally, we will study costs of the following form:
N −1
1 γ X
J∞ (ν0 , γ) = lim sup Eν0 [ c(xt , ut )]
N →∞ N t=0
18 2 Controlled Markov Chains

Such a problem is known as an infinite horizon average cost criterion.


As before, let ΓA denote the class of admissible policies, ΓM denote the class of Markov policies, ΓS denote the class of
Stationary policies. These policies can be both randomized or deterministic. We may also denote the randomized policies
with ΓRA , ΓRM and ΓRS if randomization needs to be made explicit.
For each of the criteria above, in these notes, we will investigate existence, structural and approximation results and also
computational and numerical as well as simulation based solution methods.
In a general setting, we note the following relation

inf JN (ν0 , γ) ≤ inf JN (ν0 , γ) ≤ inf JN (ν0 , γ),


γ∈ΓA γ∈ΓM γ∈ΓS

since the sets of policies are progressively shrinking

ΓS ⊂ ΓM ⊂ ΓA .

We will show, however, that for the optimal control of a Markov chain, under mild conditions, Markov policies are always
optimal (that is there is no loss in optimality in restricting the policies to be Markov); that is, it is sufficient to consider only
Markov policies. That is,
inf JN (ν0 , γ) = inf JN (ν0 , γ)
γ∈ΓA γ∈ΓM

We will also show that, under somewhat more restrictive conditions, stationary policies are optimal (that is, there is no loss
in optimality in restricting the policies to be stationary). This will typically exclude finite horizon problems and under mild
conditions we will have that

inf Jβ (ν0 , γ) = inf Jβ (ν0 , γ), and inf J∞ (ν0 , γ) = inf J∞ (ν0 , γ),
γ∈ΓA γ∈ΓS γ∈ΓA γ∈ΓRS

where we will also see that the infimum on the right hand side can be taken among those stationary policies which are
deterministic under further conditions. Furthermore, we will show that, under some stronger conditions, inf γ∈ΓS J∞ (ν0 , γ)
is independent of the initial probability measure ν0 (or the initial condition) on x0 .
For further relations between such policies, see Chapter 5 and Chapter 7.
The last two results are computationally very important, as there are powerful computational algorithms that allow one to
find such stationary policies. We will be discussing these later in the notes.
In the rest of the notes, we will first consider further properties of Markov chains, since under a Markov control policy, the
controlled state becomes a Markov chain by Theorem 2.3.1 below. We will then get back to controlled Markov chains and
the development of optimal control policies in Chapters 5 and 7.
Further optimality criteria include sample path optimality, risk-sensitive optimality and control up to a stopping time. We
will obtain structural results for optimal policies under these criteria as well, together with analytical results.

2.3.2 Stability as a Performance Criterion

In addition to, or instead of (depending on applications), optimality, one would like to achieve stability in a stochastic sense.
Such stochastic stability may be in a variety of senses, and these will be discussed in detail in the upcoming chapters.

2.3.3 Markov Chain Induced by a Markov Policy

Theorem 2.3.1 Let the control policy be randomized Markov. Then, the controlled Markov chain induces an X-valued
Markov chain, that is, the state process itself becomes a Markov chain:

Pxγ0 (xt+1 ∈ B|xt = bt , xt−1 = bt−1 , . . . , x0 = b0 ) = Qγt (xt+1 ∈ B|xt = bt ), B ∈ B(X), t ≥ 1,


2.4 Partially Observed Models and Reduction to a Fully Observed Model 19

for P almost every realization of the past variables bt , · · · , b0 , where Qγt is a possibly time-dependent stochastic ker-
nel defining a Markov chain. If the control policy is a stationary policy, then the induced Markov chain {xt } is time-
homogenous; that is, the transition kernel Qγt for the induced Markov chain does not depend on time.

Proof. We will consider the case where U is countable, the uncountable case follows similarly. Let B ∈ B(X). It follows
that,

Pxγ0 (xt+1 ∈ B|xt = bt , xt−1 = bt−1 , . . . , x0 = b0 )


= Pxγ0 (xt+1 ∈ B, ut ∈ U|xt = bt , xt−1 = bt−1 , . . . , x0 = b0 )
= Pxγ0 (∪u∈U {xt+1 ∈ B, ut = u}|xt = bt , xt−1 = bt−1 , . . . , x0 = b0 )
X
= Pxγ0 (xt+1 ∈ B, ut = u|xt = bt , xt−1 = bt−1 , . . . , x0 = b0 )
u∈U
X
= Pxγ0 (xt+1 ∈ B|ut = u, xt = bt , xt−1 = bt−1 , . . . , x0 = b0 )Pxγ0 (ut = u|xt = bt , xt−1 = bt−1 , . . . , x0 = b0 )
u∈U
X
= T (xt+1 ∈ B|ut = u, xt = bt )γt (ut = u|xt = bt )
u∈U
X
= Qγt (xt+1 ∈ B, ut = u|xt = bt )
u∈U
= Qγt (xt+1 ∈ B|xt = bt ) (2.3)

Here, Qγt is a conditional probability measure defined with Qγt (xt+1 ∈ B, ut = u|xt = bt ) := T (xt+1 ∈ B|ut = u, xt =
bt )γt (ut = u|xt = bt ). The essential issue here is that the control only depends on xt , and since xt+1 depends stochastically
only on xt and ut (being a controlled Markov chain), the desired result follows. If γt (ut |xt = bt ) = γ(ut |xt = bt ), that is,
γt = γ for all t values so that the policy is stationary, the resulting chain satisfies

Pxγ0 (xt+1 ∈ B|xt , xt−1 , . . . , x0 ) = Qγ (xt+1 ∈ B|xt ),

for some Qγ . Thus, the transition kernel does not depend on time and the chain is time-homogenous. ⋄

2.4 Partially Observed Models and Reduction to a Fully Observed Model

Consider a partially observable stochastic control problem with the following dynamics.

xt+1 = f (xt , ut , wt ), yt = g(xt , vt ).

Here, xt is the X-valued state, ut is the U-valued the control, yt is the Y-valued observation (measurement) process.
Furthermore, (wt , vt ) are i.i.d noise processes and {wt } is independent of {vt }. The controller only has causal access to
{yt }.
As noted, yt denotes an observation variable taking values in Y, a subset of Rn in the context of this review. The controller
only has causal access to the second component {yt } of the process: A deterministic admissible control policy γ is a
sequence of functions {γt } so that ut = γt (y[0,t] ; u[0,t−1] ).
We will see in Chapter 6 that one could transform a partially observable Markov Decision Problem to a Fully Observed
Markov Decision Problem via an enlargement of the state space.
Thus, the fully observed Markov Decision Model we will consider is sufficiently rich to be applicable to a large class of
controlled stochastic systems. Partially Observable Markov Decision Problems, also known as POMDPs, will be studied
in detail in Chapter 6.
20 2 Controlled Markov Chains

2.5 Decentralized Stochastic Control

We will consider situations in which there are multiple decision makers acting on a system under a variety of information
structures. These will be studied extensively in Chapter 10.

2.6 Controlled Continuous-Time Stochastic Systems

We will also study setups where the time index is a continuum. We will cover this material in Chapter 11.

2.7 Numerical Methods, Reinforcement Learning, and Robustness to Incorrect Models

While we will extensively study analytical methods to arrive at solutions, for many problems it is more convenient to
consider numerical methods or stochastic learning methods. For some applications, this may be the only option, e.g. when
a model is not known apriori. These will be studied in detail in Chapters 8 and 9.
A good control design must be robust to perturbations in the model. This brings the question of continuity and robustness
of optimal costs and optimal controls to perturbations in a model, where topological questions on model regularity are to
be studied in detail. These also, as a special case, cover finite model approximations of systems with uncountable state and
action spaces. These are studied in Chapter 12.

2.8 Bibliographic Notes

We are thankful to Prof. Eugene Feinberg on some historical remarks regarding stochastic realization and pointing out to
Aumann’s lemma: [21, Lemma F], and that this result may have also been due to Girsanov.

2.9 Exercises

Exercise 2.9.1 a) Let f be an arbitrary measurable function from R × R × R → R. Show that a controlled stochastic
process defined with
xt+1 = f (xt , ut , wt ),
with {wt } an independent and identically distributed noise sequence is a controlled Markov chain.
b) Study Lemma 3.1 and Corollary 3.1 of [56].

Exercise 2.9.2 A common example in mathematical finance applications is the portfolio selection problem where a con-
troller (investor) would like to optimally allocate his wealth between a stochastic stock market and a market with a guar-
anteed income : Consider a stock with an i.i.d. random return σt and a bank account with fixed interest rate r > 0. These
are modeled by:
Xt+1 = Xt ut (1 + σt ) + Xt (1 − ut )(1 + r), X0 = 1
and
Xt+1 = Xt (1 + r + ut (σt − r))
Here, ut ∈ [0, 1] denotes the proportion of the money that the investor invests in the stock market. Suppose that the goal is
to maximize E[log(XT )]. Then, we can write:
2.9 Exercises 21
−1
TY T −1
Xk+1 X
log(XT ) = log( )= log((1 + r + ut (σt − r))) (2.4)
Xk
k=0 k=0

Formulate the problem as an optimal stochastic control problem by clearly identifying the state and the control action
spaces, the information available at the controller, the transition kernel, and a cost functional mapping the actions and
states to R.

Exercise 2.9.3 Consider an inventory-production system given by

xt+1 = xt + ut − wt ,

where xt is R-valued, with the one-stage cost

c(xt , ut , wt ) = but + h max(0, xt + ut − wt ) + p max(0, wt − xt − ut )

Here, b is the unit production cost, h is the unit holding (storage) cost and p is the unit shortage cost; here we take p > b.
At any given time, the decision maker can take ut ∈ R+ . The demand variable wt ∼ µ is a R+ -valued i.i.d. process,
independent of x0 , with a finite mean where µ is assumed to admit a probability density function. The goal is to minimize
T
X −1
J(x, γ) = Exγ [ c(xt , ut , wt )]
t=0

The controller at time t has access to It = {xs , us , s ≤ t − 1} ∪ {xt }.


Formulate the problem as an optimal stochastic control problem by clearly identifying the state, the control action spaces,
the information available at the controller, the transition kernel and a cost functional mapping the actions and states to R.

Exercise 2.9.4 A fishery manager annually has xt units of fish and sells ut xt of these where ut ∈ [0, 1]. With the remaining
ones, the next year’s production is given by the following model

xt+1 = wt xt (1 − ut ) + wt ,

with x0 is given and wt is an independent, identically distributed sequence of random variables and wt ≥ 0 for all t and
therefore E[wt ] = w̃ ≥ 0.
The goal is to maximize the profit over the time horizon 0 ≤ t ≤ T − 1. At time T , he sells all of the fish.
Formulate the problem as an optimal stochastic control problem by clearly identifying the state, the control actions, the
information available at the controller, the transition kernel and a cost functional mapping the actions and states to R.

Exercise 2.9.5 An investor’s wealth dynamics is given by the following:

xt+1 = ut wt ,

where {wt } is an i.i.d. R+ -valued stochastic process with E[wt ] = 1. The investor has access to the past and current
wealth information and his previous actions. The goal is to maximize:
T −1
X √
J(x0 , γ) = Exγ0 [ xt − ut ].
t=0

The investor’s action set for any given x is: U(x) = [0, x].
Formulate the problem as an optimal stochastic control problem by clearly identifying the state, the control action spaces,
the information available at the controller, the transition kernel and a cost functional mapping the actions and states to R.
22 2 Controlled Markov Chains

Exercise 2.9.6 Consider an unemployed person who will have to work for years t = 1, 2, ..., 10 if she takes a job at any
given t.
Suppose that each year in which she remains unemployed; she may be offered a good job that pays 10 dollars per year
(with probability 1/4); she may be offered a bad job that pays 4 dollars per year (with probability 1/4); or she may not be
offered a job (with probability 1/2). These events of job offers are independent from year to year (that is the job market is
represented by an independent sequence of random variables for every year).
Once she accepts a job, she will remain in that job for the rest of the ten years. That is, for example, she cannot switch from
the bad job to the good job.
Suppose the goal is maximize the expected total earnings in ten years, starting from year 1 up to year 10 (including year
10).
State the problem as a Markov Decision Problem, identify the state space, the action space and the transition kernel.

Exercise 2.9.7 (Zero-Delay Source Coding) Let {xt }t≥0 be an X-valued discrete-time Markov process where X can be a
finite set or Rn . Let there be an encoder which encodes (quantizes) the source samples and transmits the encoded versions
to a receiver over a discrete noiseless channel with input and output alphabet M := {1, 2, . . . , M }, where M is a positive
integer.
The encoder policy η is a sequence of functions {ηt }t≥0 with

ηt : Mt × (X)t+1 ∋ (q[0,t−1] , x[0,t] ) 7→ qt ∈ M.

A zero-delay receiver policy is a sequence of functions γ = {γt }t≥0 of type

γt : Mt+1 ∋ q[0,t] 7→ ut ∈ U.

For the finite horizon setting the goal is to minimize the average cumulative cost (distortion)
 T −1 
T 1 X
J (π0 , η, γ) := Eπη,γ c0 (xt , ut ) , (2.5)
0
T t=0

for some T ≥ 1, where c0 : X × U → R is a nonnegative cost (distortion) function, and Eπη,γ


0
denotes expectation with
initial distribution π0 for x0 and under the quantization policy η and receiver policy γ.
Express this problem as a controlled Markov chain problem. Later on, we will provide further refinements. There is a rich
history behind this problem, see e.g., [330], [321], [305] and [335, 343].

Exercise 2.9.8 Suppose that there are two decision makers DM1 and DM2 . Suppose that the information available to to
DM1 is a random variable Y 1 and the information available to DM2 is Y 2 , where these random variables are defined on
a probability space (Ω, F, P ). Suppose that for i = 1, 2, Y i is Yi -valued and these are standard Borel spaces. Let X be a
X-valued random variable defined on the same probability space where X is also a standard Borel space.
Suppose that the sigma-field generated by Y 1 is a subset of the sigma-field generated by Y 2 , that is σ(Y 1 ) ⊂ σ(Y 2 ). That
is, the information contained in Y 1 is a subset of the information contained in Y 2 (Recall here that the σ-field generated
by a random variable Y is the smallest σ-field over Ω on which Y is measurable).
Further, suppose that the decision makers wish to minimize the following cost function:

E[c(X, U )],

where c : X × U → R+ is a measurable cost function. Let, for i = 1, 2, U i = γ i (Y i ) be generated by a measurable


function γ i on the sigma-field generated by the random variable Y i . Let Γ i denote the space of all such functions (which
we will refer to as policies).
Prove that
2.9 Exercises 23

inf E[c(X, U 1 )] ≥ 2inf 2 E[c(X, U 2 )].


γ 1 ∈Γ 1 γ ∈Γ

Hint: Make the argument that every policy u1 = γ 1 (Y 1 ) can be expressed as u2 = γ 2 (Y 2 ) for some γ 2 ∈ Γ 2 ; see Exercise
1.6.6.
3

Classification of Markov Chains

3.1 Countable State Space Markov Chains

In this section, we study Markov chains where the Markovian state takes values in a finite or a countably infinite set X.
In the following, we will consider (Ω, F, P) to be the probability space on which all of the random variables are defined
(later on, when a particular notational distinction is not needed, we will replace the notation P with P as the probability
measure on the events to be considered).
We assume that ν0 is an initial distribution for the Markov chain, so that P(x0 ∈ ·) = ν0 (·) (also denoted with x0 ∼ ν0 ).
The process Φ = {x0 , x1 , . . . , xn , . . . } is a (time-homogeneous) Markov chain with its probability law (or the probability
measure induced on the sequence space) satisfying, for n ∈ Z+ :

Pν0 (x0 = a0 , x1 = a1 , x2 = a2 , . . . , xn = an )
:= ν0 (x0 = a0 )P(x1 = a1 |x0 = a0 )P(x2 = a2 |x1 = a1 ) . . . P(xn = an |xn−1 = an−1 ) (3.1)

If the initial condition is known to be a fixed state a0 ∈ X, we use Pa0 (· · · ) in place of Pδa0 (· · · ). We could represent the
probabilistic evolution in terms of a matrix:

P (i, j) := P(xt+1 = j|xt = i) ≥ 0, i, j ∈ X.

Here PP (·, ·) is a probability transition kernel, that is for every i ∈ X, P (i, .) is a probability measure on X, in particular
with j P (i, j) = 1 for every i. Let P be the |X| × |X| matrix with entries given with P (i, j) ≥ 0. Such a matrix P is
called a stochastic matrix.
The initial condition probability and the transition kernel uniquely identify the probability measure on the product space
XN , by the extension theorems presented in Chapter 1.
Let πk (i) = P(xk = i) for i ∈ X, and for all k ∈ Z+ . Let πk = [πk (i), i ∈ X]. It follows that
X X X
π1 (j) = P(x1 = j) = P(x1 = j, x0 = i) = P(x1 = j|x0 = i)P(x0 = i) = π0 (i)P (i, j)
i∈X i∈X i∈X

and with P denoting the transition matrix given with P (i, j) as defined above, by a similar reasoning,

πk+1 = πk P, k ∈ Z+ (3.2)

Note here that we represent πk as a row vector. By induction, we could verify that for k ∈ N:
X
P k (i, j) := P (xt+k = j|xt = i) = P (i, m)P k−1 (m, j)
m∈X
26 3 Classification of Markov Chains

We will see that whether the sequence {πk , k ∈ Z+ } admits a limit and the dependence properties of this limit on π0 have
significant implications on the characterization of Markov chains, and later, in stabilization and optimization of controlled
Markov chains.
In the following, we characterize Markov Chains based on transience, recurrence and communication properties. We then
consider the problem of the existence of an invariant probability measure. Later, we will extend the analysis to uncountable
space Markov chains.

Communication

If there exists an integer k ∈ N such that P(xt+k = j|xt = i) = P k (i, j) > 0, and an integer l ∈ N such that
P (xt+l = i|xt = j) = P l (j, i) > 0 then state i communicates with state j.
A set C ⊂ X is said to be communicating if every two elements (states) of C communicate with each other.
If every member of the set communicates with every other member, such a chain is said to be irreducible.
The period of a state i ∈ X is defined to be the greatest common divisor of {k > 0 : P k (i, i) > 0}.
A Markov chain is called aperiodic if the period of all states is 1.

Absorbing Set

A set C is called absorbing if P (i, C) = 1 for all i ∈ C. That is, if the state is in C, then the state cannot get out of C.
The Markov chain is irreducible if the smallest absorbing set is the entire X itself.
The Markov chain is indecomposable if X does not contain two disjoint absorbing sets.

Occupation, Hitting and Stopping Times

For any set A ⊂ X, the occupation time ηA is the number of visits of {xt } to set A:

X
ηA = 1{xt ∈A} ,
t=0

where 1E denotes the indicator function for an event E, that is, it takes the value 1 when E takes place, and is otherwise 0.

Remark 3.1. Another common notation for the indicator function is the following: Let A be an event (a subset of some
σ-field). Then 1A (x) = 1 if x ∈ A and 0 otherwise.

Let A ⊂ X. Define
τA := min{k > 0 : xk ∈ A},
to be the first time that the state visits A; we call this the return time to set A. We also define a very similar notion, called a
hitting time:
σA = min{k ≥ 0 : xk ∈ A}.
The variable τA defined above is an example for stopping times:

Definition 3.1.1 A Z+ ∪ {∞}-valued random variable τ is a stopping time (with respect to the σ-field generated by the
process {x0 , x1 , · · · }), if for all n ∈ Z+ , the event {τ = n} ∈ σ(x0 , x1 , x2 , . . . , xn ), that is the event is in the sigma-field
generated by the random variables up to time n.
3.1 Countable State Space Markov Chains 27

Any realistic decision takes place at a time which is a stopping time. Consider an optimal investment problem: if an investor
claims to stop investing (e.g., purchasing houses) when the investment (value of the housing market) is at its local peak, the
decision instant could not be a stopping-time in general: this peak-time is not a stopping time because to find out whether
the investment value is at its peak, the next realization should be known, and this information is not available up to any
given time in a causal fashion for a non-trivial (i.e., non-deterministic) stochastic process.
One important property of Markov chains is the strong Markov property. This says the following: If we sample a Markov
chain according to a stopping time rule, the sampled Markov chain starts from the sampled instant as a Markov chain with
the same transition probabilities as if the sampling instant is time 0:

Proposition 3.1.1 For a (time-homogenous) Markov chain with a countable state space X, the strong Markov property
holds: that is, if τ is a stopping time with P (τ < ∞) = 1, then almost surely for any m ∈ N:

P (xτ +m = a|xτ = b0 , xτ −1 = b1 , . . . ) = P (xτ +m = a|xτ = b0 ) = P m (b0 , a).

Proof. We consider m = 1; for larger m, the result follows from identical steps. For an event with {xτ = b0 , xτ −1 =
b1 , . . . } with P (xτ = b0 , xτ −1 = b1 , . . . ) > 0, we have that

P (xτ +1 = a|xτ = b0 , xτ −1 = b1 , . . . )
P (xτ +1 = a, xτ = b0 , xτ −1 = b1 , . . . )
=
P (xτ = b0 , xτ −1 = b1 , . . . )
P∞
P (τ = k, xτ +1 = a, xτ = b0 , xτ −1 = b1 , . . . )
= k=0 (3.3)
P (xτ = b0 , xτ −1 = b1 , . . . )
P∞
P (xk+1 = a|τ = k, xk = b0 , xk−1 = b1 , . . . )P (τ = k, xk = b0 , xk−1 = b1 , . . . )
= k=0
P (xτ = b0 , xτ −1 = b1 , . . . )
P∞
P (xk+1 = a|xk = b0 , xk−1 = b1 , . . . )P (τ = k, xk = b0 , xk−1 = b1 , . . . )
= k=0 (3.4)
P (xτ = b0 , xτ −1 = b1 , . . . )
P∞
P (xk+1 = a|xk = b0 )P (τ = k, xk = b0 , xk−1 = b1 , . . . )
= k=0
P (xτ = b0 , xτ −1 = b1 , . . . )
P∞
P (τ = k, xτ = b0 , xτ −1 = b1 , . . . )
= P (b0 , a) k=0
P (xτ = b0 , xτ −1 = b1 , . . . )
P (xτ = b0 , xτ −1 = b1 , . . . )
= P (b0 , a)
P (xτ = b0 , xτ −1 = b1 , . . . )
= P (b0 , a) (3.5)

Note that the assumption P (τ < ∞) = 1 is critically used in the proof in (3.3). In (3.4), we use the fact that τ is a stopping
time. ⋄

3.1.1 Recurrence and transience

Let us define
X∞ ∞
X X∞
U (x, A) := E[ 1(xt ∈A) |xo = x] = P t (x, A) =: Ex [ 1(xt ∈A)
t=1 t=1 t=1

and define
L(x, A) := P (τA < ∞|x0 = x) =: Px (τA < ∞),
which is the probability of the chain visiting set A, once the process starts at state x.

Definition 3.1.2 (i) A set A ⊂ X is recurrent if the Markov chain visits A infinitely often in expectation, when the process
starts in A:
28 3 Classification of Markov Chains

Ex [ηA ] = ∞, ∀x ∈ A (3.6)

(ii) A state α ∈ X is transient if

U (α, α) = Eα [ηα ] < ∞. (3.7)

(iii)A set A ⊂ X is positive recurrent if

Ex [τA ] < ∞, ∀x ∈ A.

In particular, if a state α ∈ X is not recurrent, it is transient.


P∞
Equation (3.7) can also be written as i=1 P i (α, α) < ∞, which in turn is implied by

Pi (τi < ∞) < 1,

as we will show further below. The reader should connect the above with the strong Markov property: once the process hits
a state, it starts from the state as if it is time 0 (regardless of the the past); the process recurs itself.
There is another important notion of recurrence, called Harris recurrence:

Definition 3.1.3 (i) A set A is Harris recurrent if Px (ηA = ∞) = 1 for all x ∈ A.


(ii) An irreducible Markov chain is Harris recurrent if

Px (ηA = ∞) = 1, ∀x ∈ X, A ⊂ X.

Let τi (1) := τi and for i ≥ 1,


τi (k + 1) = min{n > τi (k) : xn = i}

We have the following result whose proof, which builds on continuity of probability (Theorem B.1.2), is presented later in
the chapter in a more general context in Theorem 3.2.1.

Theorem 3.1.1 The condition Pi (τi < ∞) = 1 is equivalent to the condition Pi (ηi = ∞) = 1.

One can verify that (3.7) is equivalent to L(i, i) < 1.

Theorem 3.1.2 If Pi (τi < ∞) < 1, then Ei [ηi ] < ∞ and thus the state i ∈ X is transient.

To show this, it suffices to first verify the relation

Pi (τi (k) < ∞) = Pi (τi (k − 1) < ∞)Pi (τi (1) < ∞),
P∞
and then use the equality E[η] = k=1 P (η ≥ k).
We will investigate the Harris recurrence property further while studying uncountable state space Markov chains, however
one needs to note that even for countable state space chains Harris recurrence is stronger than recurrence as we make
explicit next.

Remark 3.2. Harris recurrence is stronger than recurrence. In one, an expectation is considered; in the other, a probability
is considered. Consider the following example: Let X = N, P (1, 1) = 1 and for x > 1: P (x, x + 1) = 1 − 1/x2 and
P (x, 1) = 1/x2 . Then, for x ≥ 2 (see Exercise 3.5.7):
Y
Px (τ1 = ∞) = (1 − 1/t2 ) > 0.
t≥x,t∈N
3.1 Countable State Space Markov Chains 29

Thus, the set {1, 2} is not Harris recurrent, but it is recurrent.

3.1.2 Stability and invariant measures

Stability is an important concept, but it has different meanings in different contexts. This notion will be made more precise
in the following chapter. Nonetheless, perhaps the weakest form of stochastic stability in the context of these notes is the
existence of an invariant probability measure.
Recall from (3.2) that the occupation probabilities satisfy the recursions:

π1 = π0 P

And for t > 1:


πt+1 = πt P = π0 P t+1
One important property of Markov chains is whether the above iteration leads to a fixed point. Such a fixed point π is called
an invariant probability measure. Thus, a probability measure in a countable state Markov chain is invariant if

π = πP

This is equivalent to X
π(j) = π(i)P (i, j), ∀j ∈ X
i∈X

We note that, if such a π exists, it must be written in terms of π = π0 limt→∞ P t , for some π0 . Clearly, π0 can be π itself,
but often π0 can be any initial probability measure P under irreducibility/aperiodicity conditions (where aperiodicity can be
T −1
relaxed if convergence of the averages limt→∞ T1 t=0 π0 P t is considered) which will be discussed further. Invariant
probability measures are especially important in stochastic control, due to ergodicity theorems (which show that temporal
averages
PT −1 converge to statistical averages with probability 1), as we will discuss later in the chapter. Finally, how fast
1 t
T t=0 π 0 P converges to invariance is another very important question to be studied.

3.1.3 Invariant measures via an occupational characterization

The following is one of the most consequential results in this chapter.

Theorem 3.1.3 For a Markov chain, if there exists an element i such that Ei [τi ] < ∞; the following is an invariant
probability measure:
 Pτi −1 
k=0 1{xk =j}
µ(j) = E x0 = i , j∈X
Ei [τi ]

Proof. We will show for every j ∈ X that


 Pτi −1  X  Pτi −1 
k=0 1{xk =j} k=0 1{xk =s}
E x0 = i = P (s, j)E x0 = i ,
E[τi ] E[τi ]
s∈X

which establishes the desired result1 . Note that E[1{Xt+1 =j} ] = P (Xt+1 = j) and P (s, j) = E[1{Xk+1 =j} |Xk = s] =
E[1{Xk+1 =j} |Xk ](ω) with Xk (ω) = s. Hence,
 Pτi −1 
X
k=0 1{Xk =s}
P (s, j)E X0 = i
Ei [τi ]
s∈X
1
In the following, to make the random nature of xk terms explicit, we will use capital letters Xk to emphasize randomness. In the
notes, we will occasionally follow this route, since for conditional expectations, often it is very crucial to distinguish between random
variables and their realizations.
30 3 Classification of Markov Chains
 Pτi −1 P
P (s, j)1{Xk =s}

k=0 s∈X
=E X0 = i
Ei [τi ]
   
Pτi −1 P
E k=0 s∈X 1{Xk =s} E[1{Xk+1 =j} |Xk = s] x0 = i
=
Ei [τi ]
   
Pτi −1 P
E k=0 s∈X 1{Xk =s} E[1{Xk+1 =j} |Xk ] x0 = i
= (3.8)
Ei [τi ]
 
Pτi −1 P
E k=0 1
s∈X {Xk =s} E[1 {xk+1 =j} |Xk , X k−1 , X k−2 , · · · , X 0 = i] X 0 = i
= (3.9)
Ei [τi ]
 
Pτi −1 P
E k=0 s∈X E[1{Xk =s} 1{Xk+1 =j} |Xk , Xk−1 , Xk−2 , · · · , X0 = i] X0 = i
= (3.10)
Ei [τi ]
 
P∞ P
E k=0 s∈X 1{k<τi } E[1{Xk =s} 1{Xk+1 =j} |Xk , Xk−1 , Xk−2 , · · · , X0 = i] X0 = i
=
Ei [τi ]
 
P∞ P
k=0 s∈X E E[1 1 1
{k<τi } {Xk =s} {Xk+1 =j} |X k , X k−1 , X k−2 , · · · , X 0 = i] X 0 = i
= (3.11)
Ei [τi ]
 
P∞ P
k=0 E 1{k<τi } 1 1
s∈X {Xk =s} {Xk+1 =j} X 0 = i
= (3.12)
Ei [τi ]
 
P∞
E k=0 1{k<τi } 1{Xk+1 =j} |X0 = i
=
Ei [τi ]
 
Pτi −1
E k=0 1{Xk+1 =j} |X0 = i
 Pτi −1
k=0 1{Xk+1 =j}

= = Ei
Ei [τi ] Ei [τi ]
 Pτi   Pτi −1 
k=1 1{Xk =j} k=0 1{Xk =j}
= Ei = Ei
Ei [τi ] Ei [τi ]
= µ(j),

where we use the fact that the total number of visits to a given set does not change whether we include either t = 0 or
τi , since X0 = Xτi = i. Here, (3.8) follows from the fact that Xk = s is specified so that 1{Xk =s} E[1{Xk+1 =j} |Xk ] =
1{Xk =s} E[1{Xk+1 =j} |Xk = s], (3.9) follows from the fact that the process is a Markov chain, (3.10) and (3.11) follow
from the properties of conditional expectation and that τi is a stopping time (we will discuss such properties in Chapter 4),
and (3.12) follows from the law of the iterated expectations, see Theorem 4.1.3. In the above (3.8) follows from the fact
that 1{Xk =s} E[1{Xk+1 =j} |Xk ] = E[1{Xk =s} 1{Xk+1 =j} |Xk ].
Finally, observe that if Ei [τi ] < ∞, then the above measure indeed is a probability measure, as it follows that

X  τi −1 1{Xk =j}
P 
X
k=0
µ(j) = E X0 = i = 1.
j
E[τi ]

This concludes the proof. ⋄

Theorem 3.1.4 Every finite state space Markov chain admits an invariant probability measure.
3.1 Countable State Space Markov Chains 31

A common proof technique on the existence of invariant probability measures for finite state Markov chains builds on
an important result called the Perron-Frobenius Theorem. However, we will present a more comprehensive result in the
context of general space Markov chains later in Theorem 3.3.1.

Theorem 3.1.5 For an irreducible Markov chain with countable X, there can be at most one invariant probability measure.

Proof. Let π(i) and π ′ (i) be two different invariant probability measures. Define D := {i : π(i) > π ′ (i)}. Then,
X X
π(D) = π(i)P (i, D) + π(i)P (i, D)
i∈D i∈D
/
X X
π ′ (D) = π ′ (i)P (i, D) + π ′ (i)P (i, D)
i∈D i∈D
/

implies that X X
π(D) − π ′ (D) = (π(i) − π ′ (i))P (i, D) + (π(i) − π ′ (i))P (i, D)
i∈D i∈D
/

and thus X X
(π(i) − π ′ (i))(1 − P (i, D)) = (π(i) − π ′ (i))P (i, D)
i∈D i∈D
/

The first term is strictly positive (since P (i, D) = 1 cannot hold for all i ∈ D due to irreducibility, for otherwise D would
be absorbing). The second term is not positive, hence a contradiction. ⋄

Remark 3.3. One can see that for any a ∈ X with π(a) > 0, it must be that Ea [τa ] < ∞. The reason will be evident once
we study Theorem 3.2.7 and consider the relation:
a −1
τX
1 = π(X) = π(a)Ea [ 1{xk ∈X} ] = π(a)Ea [τa ]
k=0

An implication of the above is the following very important result, which is a special case of Kac’s lemma (see also
Theorem

Theorem 3.1.6 (Kac’s Lemma) Let {xt } be irreducible and π be its (unique) invariant probability measure. Then, for all
i ∈ X with π(i) > 0,
1
π(i) = , i ∈ X.
Ei [τi ]

Remark 3.4. Consider the random walk on Z given with the transition kernel P (x, x + 1) = P (x, x − 1) = 21 for z ∈ Z.
In this case, we have that for every i ∈ Z, Ei [τi ] = ∞, and hence there does not exist an invariant probability measure.
But, it has an invariant measure defined with: µ({i}) = K, i ∈ Z, for an arbitrary (fixed) K ∈ R. That Ei [τi ] = ∞
can be established through the following reasoning: if there were an invariant probability measure, this would be unique
by irreducibility and also for every state i, the measure Ei1[τi ] would take the same value. But the sum of these (countably
infinitely many) identical values would need to be 1, leading to a contradiction. Then, Ei [τi ] cannot be finite for any i.

3.1.4 Rates of convergence to invariant measures and Dobrushin’s ergodic coefficient

Consider the iteration πt+1 = πt P , with a given π0 . We would like to know when this iteration converges to a limit and
how fast this convergence is. Here, the reader is referred to Appendix A for a review of vector and function spaces.
A map T from one complete normed linear (that is, a Banach) space X to itself is called a contraction if for some 0 ≤ ρ < 1

∥T (x) − T (y)∥ ≤ ρ∥x − y∥, ∀x, y ∈ X.


32 3 Classification of Markov Chains

Theorem 3.1.7 A contraction map T in a Banach space has a unique fixed point x∗ with x∗ = T (x∗ ). Furthermore, the
iterates xn+1 = T (xn ), for any given x0 , converge to x∗ geometrically fast in the sense that ∥xn − x∗ ∥ ≤ Lx0 ρn for some
Lx0 < ∞.

Proof. {T n (x)} forms a Cauchy sequence: First note that, ∥T k (x) − T k−1 (x)∥ ≤ ∥T (T k−1 (x)) − T (T k−2 (x))∥ ≤
ρ∥T k (x) − T k−1 (x)∥ ≤ · · · ≤ ρk−1 ∥T (x) − x)∥. Then,
n n
X X 1
∥T n (x) − x∥ ≤ ∥T k (x) − T k−1 (x)∥ ≤ ρk−1 ∥T (x) − x)∥ ≤ ∥T (x) − x)∥
1−ρ
k=1 k=1

implying that
1
∥T n (x)∥ ≤ ∥x∥ + ∥T n (x) − x∥ ≤ ∥x∥ + ∥T (x) − x)∥ =: M (x)
1−ρ
uniformly over all n. Now, for every n, m ≥ N , we have that ∥T n (x) − T m (x)∥ ≤ ρN ∥T n−N (x) − T m−N (x)∥ ≤
2M (x)ρN . This implies that the sequence is Cauchy. By completeness, the Cauchy sequence has a limit, x∗ . For
uniqueness, suppose that there are two (different) fixed points with u = T (u) and v = T (v). Then ∥u − v∥ =
∥T (u) − T (v)∥ ≤ ρ∥u − v∥ < ∥u − v∥, a contradiction. Thus, u = v. The rate of convergence follows by writing
∥xn − x∗ ∥ = ∥T n (x0 ) − T n (x∗ )∥ ≤ ρn ∥x0 − x∗ ∥. ⋄
Contraction Mapping via Dobrushin’s Ergodic Coefficient Consider a countable state Markov Chain with one-step
transition kernel P . Define the Dobrushin coefficient as
X 
δ(P ) = min min(P (i, j), P (k, j)) (3.13)
i,k
j∈X

Observe that for two scalars a, b


|a − b| = a + b − 2 min(a, b).
Let us define for a vector v the l1 norm: X
||v||1 = |vi |.
i∈X

The set of all countable index real-valued vectors (that is functions which map Z → R) with a finite l1 norm

{v : ||v||1 < ∞}

is a complete normed linear space, and as such, is a Banach space. With these observations, we state the following:

Theorem 3.1.8 [Dobrushin] [105] For any two probability measures π, π ′ , it follows that

||πP − π ′ P ||1 ≤ (1 − δ(P ))||π − π ′ ||1 .

Accordingly, the sequence of iterates πn+1 = T (πn ) := πn P , for any given π0 , converges to invariance geometrically
fast.

Proof. Let ψ(i) = π(i) − min(π(i), π ′ (i)) for all i ∈ X. Further, let ψ ′ (i) = π ′ (i) − min(π(i), π ′ (i)). Since
X X X
0= π(i) − π ′ (i) = π(i) − π ′ (i) + π(i) − π ′ (i)
i i:π(i)>π ′ (i) i:π ′ (i)>π(i)

we have that ||ψ||1 = ||ψ ′ ||1 , and since


X X X
|π(i) − π ′ (i)| = ψ(i) + ψ ′ (i)
i i:π(i)>π ′ (i) i:π ′ (i)>π(i)

we have that
3.1 Countable State Space Markov Chains 33
X
|π(i) − π ′ (i)| = ||ψ||1 + ||ψ ′ ||1
i

and thus
||π − π ′ ||1 = ||ψ − ψ ′ ||1 = 2||ψ||1 = 2||ψ ′ ||1
Now,

||πP − π ′ P ||1 = ||ψP − ψ ′ P ||1


X X X
= | ψ(i)P (i, j) − ψ ′ (k)P (k, j)|
j i k
1 X XX
= | ψ(i)ψ ′ (k)P (i, j) − ψ(i)ψ ′ (k)P (k, j)| (3.14)
||ψ ′ ||1 j i k
1 XXX
≤ ψ(i)ψ ′ (k)|P (i, j) − P (k, j)| (3.15)
||ψ ′ ||1 j i k
1 XX X
= ′
ψ(i)ψ ′ (k) |P (i, j) − P (k, j)|
||ψ ||1 i j
k
X 
1 XX ′
= |ψ(i)||ψ (k)| P (i, j) + P (k, j) − 2 min(P (i, j), P (k, j)) (3.16)
||ψ ′ ||1 i j
k
1 XX
≤ |ψ(i)||ψ ′ (k)|(2 − 2δ(P )) (3.17)
||ψ ′ ||1 i
k
= ||ψ ′ ||1 (2 − 2δ(P )) (3.18)
= ||π − π ′ ||1 (1 − δ(P ))

In the above, (3.14) follows from adding terms in the summation, (3.15) from taking the norm inside, (3.16) follows from
the relation ||a − b|| = a + b − 2 min(a, b), (3.17) from the definition of δ(P ) and finally (3.18) follows from the l1 norms
of ψ, ψ ′ .
Thus, the map πP : π ∈ P(X) 7→ πP ∈ P(X), where P(X) is the set of probability measures on X viewed as a subset of
l1 (X; R), is a contraction mapping if δ(P ) > 0. As a result, the sequence {π0 P n , n ∈ Z+ } is Cauchy by Theorem 3.1.7,
and as every Cauchy sequence in a Banach space has a limit, so does this process. We emphasize that the set of probability
measures is not a linear space, but viewed as a closed subset of l1 (X; R), the sequence will have a limit. Since πP is also a
probability measure for every π ∈ P(X), the limit must also be a probability measure. The limit is the invariant probability
measure. ⋄
It should be emphasized that Dobrushin’s theorem tells us how fast the sequence of probability measures {π0 P n } converges
to the invariant probability measure π for an arbitrary π0 : since πP n = π, we have that

∥π0 P n − π∥1 = ∥π0 P n − πP n ∥1 ≤ (1 − δ(P ))n ∥π0 − π∥1 ≤ (1 − δ(P ))n , n ∈ Z+

3.1.5 Ergodic theorem for countable state space chains

In Exercise 4.5.11, we will prove the ergodic theorem: let {xt } be a Harris recurrent Markov chain with an invariant
probability measure µ (such a process is called positive Harris recurrent, as we will define in the next section). We then
have that for every fixed initial state, almost surely
T
1X X
lim f (xt ) = f (i)µ(i)
T →∞ T
t=1 i
P
for bounded f (if f is not bounded, we require that i |f (i)|µ(i) < ∞). This is a very important theorem, as this property
is what establishes an important connection with average cost stochastic control. Under a stationary control policy leading
34 3 Classification of Markov Chains

to a unique invariant probability measure µ on the state and control process (which is a Markov chain), with a bounded
function c it follows that almost surely,
T
1X X
lim c(xt , ut ) = c(x, u)µ(x, u).
T →∞ T
t=1 x,u

3.2 Uncountable Standard Borel State Spaces

We now extend the discussion above to the uncountable state space setting. We will consider state spaces that are standard
Borel; as noted earlier, these are Borel subsets of complete, separable and metric spaces. We note again that the spaces that
are complete, separable and metric are also called Polish metric spaces.
Let {xt , t ∈ Z+ } be a Markov chain with a Polish state space X, and defined on a probability space (Ω, F, P), where Ω
is the sample space, F a sigma field of subsets of Ω, and P a probability measure. Let P (x, D) := P(xt+1 ∈ D|xt = x)
denote the transition probability from x to D, that is the probability of the event {xt+1 ∈ D} given that xt = x.
We could compute P (xt+k ∈ D|xt = x) inductively as follows:
Z Z Z
P(xt+k ∈ D|xt = x) = · · · P (xt , dxt+1 ) . . . P (xt+k−2 , dxt+k−1 )P (xt+k−1 , D)

As such, we have for all n ≥ 1, states x and Borel sets A, P n (x, A) := P(xt+n ∈ A|xt = x) = P n−1 (x, dy)P (y, A),
R
X
with P 1 (·, ·) := P (·, ·).

Definition 3.2.1 A Markov chain is µ-irreducible, if for any set B ∈ B(X) such that µ(B) > 0, and any x ∈ X, there
exists some integer n > 0 (possibly depending on B and x), such that P n (x, B) > 0, where P n (x, B) is the transition
probability in n stages, that is, P (xt+n ∈ B|xt = x).

A maximal irreducibility measure ψ is an irreducibility measure such that for all other irreducibility measures ϕ, we have
ψ(B) = 0 ⇒ ϕ(B) = 0 for any B ∈ B(X ) (that is, all other irreducibility measures are absolutely continuous with respect
to ψ). In the text, whenever a chain is said to be irreducible, irreducibility with respect to a maximal irreducibility measure
is implied. We also define B + (X ) = {A ∈ B(X ) : ψ(A) > 0} where ψ is a maximal irreducibility measure. A maximal
irreducibility measure ψ exists for a µ-irreducible Markov chain, for example ψ(B) = n∈Z+ 2−n P n (x, B)µ(dx) (see
P
[232, Propostion 4.2.2]).
As an example, consider the following linear system:

xt+1 = axt + wt ,

This chain is Lebesgue irreducible if wt is a Gaussian variable. The definitions for recurrence and transience follow those
in the countable state space setting:

Definition 3.2.2 A set A ∈ B(X) is called recurrent if

X∞ ∞
X
Ex [ 1xt ∈A ] = P t (x, A) = ∞, ∀x ∈ A.
t=1 t=1

A ψ−irreducible Markov chain is called recurrent if, for A with ψ(A) > 0,

X∞ ∞
X
Ex [ 1xt ∈A ] = P t (x, A) = ∞, ∀x ∈ X.
t=1 t=1

Definition 3.2.3 A set A ∈ B(X) is Harris recurrent if


3.2 Uncountable Standard Borel State Spaces 35

Px (ηA = ∞) = 1, ∀x ∈ A. (3.19)

A ψ-irreducible Markov chain is Harris recurrent if

Px (ηA = ∞) = 1, A ∈ B(X), ψ(A) > 0, ∀x ∈ X.

Theorem 3.2.1 Harris recurrence of a set A is equivalent to

Px (τA < ∞) = 1, ∀x ∈ A.

Proof. Let τA (1) be the first time the state hits A. Now, with x ∈ A, the probability of τA (2) < ∞ can be computed
recursively as

Px (τA (2) < ∞) = Px (τA (2) < ∞, τA (1) < ∞)


X
= Px (τA (2) < ∞, τA (1) = m)
m∈Z+
X
= Px (τA (2) < ∞, XτA (1) ∈ A, τA (1) = m)
m∈Z+
X Z
= Px (τA (2) < ∞, XτA (1) ∈ dy, τA (1) = m)
m∈Z+ A
X
= Px (τA (2) < ∞|XτA (1) = y, τA (1) = m)Px (XτA (1) ∈ dy, τA (1) = m)
m∈Z+
X Z
= P (τA (1) < ∞|X0 = y)Px (XτA (1) ∈ dy, τA (1) = m) (3.20)
m∈Z+ A
X Z
= Px (XτA (1) ∈ dy, τA (1) = m)
m∈Z+ A
X
= Px (XτA (1) ∈ A, τA (1) = m)
m∈Z+
X
= Px (τA (1) = m) = 1,
m∈Z+

where (3.20) uses the strong Markov property and the next equation follows from P (τA (1) < ∞|X0 = y) = 1 for y ∈ A.
By induction, for every n ∈ N

Px (τA (n + 1) < ∞) = 1 (3.21)

Now,
Px (ηA ≥ k) = Px (τA (k − 1) < ∞),
since k times visiting a set requires k times returning to a set, when the initial state x is in the set. As such,

Px (ηA ≥ k) = 1, ∀k ∈ Z+

is identically equal to 1. Define Bk = {ω ∈ Ω :Tη(ω) ≥ k}, and it follows that Bk+1 ⊂ Bk for all k ∈ N. By the

continuity of probability (see Theorem B.1.2), P ( k=1 Bk ) = limk→∞ P (Bk ), it follows that Px (ηA = ∞) = 1. The
other direction for equivalence follows from the definitions of occupation time ηA and return time τA . ⋄

Definition 3.2.4 For a Markov chain with transition probability P , a probability measure π is invariant if
Z
π(D) = P (x, D)π(dx), D ∈ B(X).
X
36 3 Classification of Markov Chains

Definition 3.2.5 If a Harris recurrent Markov chain admits an invariant probability measure, then the chain is called
positive Harris recurrent.

We will discuss the ergodic theorem for such chains further, but it may be useful to state the following:

Lemma 3.2.1 (Ergodic Theorem for Positive Harris Recurrent Markov Chains, MeynBook] For a Markov chain {Xn }n∈N
which admits at least one invariant probability measure, the following statements are equivalent:

(i) The chain is positive Harris recurrent.


(ii) There exists an invariant probability measure π of such that for all f ∈ L1 (π) and every initial distribution µ,
n Z !
1 X
P lim f (xi ) = f (x)π(dx) = 1,
n→∞ n + 1 X
i=0

where {xi , i ∈ Z+ } is Markov with x0 ∼ µ.

3.2.1 Invariant probability measures and split chains

Uncountable chains act like countable ones when there is a single atom α ⊂ X which satisfies a finite mean return property
to be discussed below.

Definition 3.2.6 A set α is called an atom if there exists a probability measure ν such that

P (x, A) = ν(A), ∀x ∈ α, ∀A ∈ B(X).

If the chain is ψ-irreducible and ψ(α) > 0, then α is called an accessible atom.

In case there is an accessible atom α, we have the following result the proof of which follows the same steps of those of
Theorem 3.1.3 and 3.1.5.

Theorem 3.2.2 For a ψ-irreducible Markov chain for which Eα [τα ] < ∞, the following is the invariant probability
measure:  
Pτα −1
Eα k=0 1{xk ∈A}
π(A) = , A ∈ B(X)
Eα [τα ]

Small Sets and Nummelin and Athreya-Ney’s Splitting Technique

In case an atom is not present, we may be able to construct an artificial atom:

Definition 3.2.7 A set A ∈ B(X) is (n-µ)-small on (X, B(X)) if for some non-trivial positive (i.e., not all sets have zero
measure) measure µ and n ∈ N

P n (x, B) ≥ µ(B), ∀x ∈ A, and B ∈ B(X)

Often, we simply say that a set is small without specifying the smallness measure µ or time index n.
The results on recurrence apply to uncountable chains with no atom provided there is a small set or a petite set (to be
discussed further below). In the following, we construct an artificial atom through what is commonly known as the splitting
technique, see [246] [247] (see also [20]).
3.2 Uncountable Standard Borel State Spaces 37

Suppose a set A is 1-small. Define a process zt = (xt , at ), where zt is ∈X ×{0, 1}-valued and {at } is i.i.d. Bernoulli with
P (at = 1) = δ. That is, we enlarge the state space and observe that zt is also Markov. When xt ∈ / A, {at }, {xt } evolve
independently from each other. When xt ∈ A, depending on the realization of at , with probability δ the state is mapped to
A × {1} and with probability 1 − δ the state is mapped to A × {0}. From A × {1}, the transition for the next time stage is
t)
given by ν(dx
δ (for all (x, 1) ∈ A × {1}), and from A × {0}, it visits the future time stage with transition probability

P (dxt+1 |xt ) − ν(dxt+1 )


,
1−δ
where δ = ν(X). That is,

ν(B)
P (xt+1∈B |(xt , at ) = (x, 1)) = , (x, 1) ∈ A × {1}
δ
P (x, B) − ν(B)
P (xt+1∈B |(xt , at ) = (x, 0)) = , (x, 0) ∈ A × {0} (3.22)
1−δ

In this case, A × {1} is an accessible atom for the extended (split) Markov chain (xt , at ), and the marginal distribution of
the original Markov process {xt } has not been altered by this construction.
The following can be established using the construction above.

Proposition 3.2.1 If
sup E[min(t > 0 : xt ∈ A)|x0 = x] < ∞
x∈A

then,
sup E[min(t > 0 : zt ∈ (A × {1}))|z0 = z] < ∞.
z∈(A×{1})

Now suppose that a set A is m-small. Then, we can construct a split chain for the sampled process xmn , n ∈ N. Note that
this sampled chain has a transition kernel as P m . We replace the discussion for the 1-small case with the sampled chain
(also known as the m-skeleton of the original chain). If one can show that the sampled chain has an invariant measure πm ,
then (see [232, Theorem 10.4.5]):
m−1 Z
1 X
π(B) := πm (dx)P k (x, B) (3.23)
m
k=0

is invariant for P . Furthermore, π is also invariant for the sampled chain with kernel P m . Hence if P m leads to a unique
invariant probability measure, π = πm .
From small to petite sets

Definition 3.2.8 [232] A set A ∈ B(X) is νR -petite on (X, B(X)) if for some distribution R on N, and some non-trivial
positive measure νR ,
X∞
P n (x, B)R(n) ≥ νR (B), ∀x ∈ A, and B ∈ B(X).
n=0

By [232, Proposition 5.5.6], if a Markov chain is ψ-irreducible and if a set C is ν-petite, then R can be taken to be a
geometric distribution aϵ (i) = (1 − ϵ)ϵi , i ∈ Z+ (with the randomly sampled chain also known as the resolvent kernel).
Another useful result to be utilized later is as follows.

Theorem 3.2.3 [232] Let {xt } be µ-irreducible and let A be ν-petite. Then, there exists a sampling distribution such that
A is ψ-petite where ψ is a maximal irreducibility measure. Furthermore, A is ψ-petite for a sampling distribution with
finite mean.
38 3 Classification of Markov Chains

Definition 3.2.9 A ψ-irreducible Markov chain is periodic with period d if there exists a partition of X = ∪di=1 Xi ∪ D so
that P (x, Xi+1 ) = 1 for all x ∈ Xi and P (x, X1 ) = 1 for all x ∈ Xd , with ψ(D) = 0. If no such d > 1 exists, the chain is
aperiodic.

Another useful result is the following.

Theorem 3.2.4 [232, Theorem 5.5.3] For an aperiodic and irreducible Markov chain {xt } every petite set is ν-small for
some appropriate ν (but now ν may not be a maximal irreducibility mesure; compare with Theorem 3.2.3).

The discussion up to (3.23) and the split chain argument applies also for an arbitrary sampling distribution K on N. Suppose
that we have
Z X 
n
πK (dx) K(n)P (x, B) = πK (B), B ∈ B(X) (3.24)
n

Then,
Z X m−1
X
π(B) := K(m) πK (dx)P k (x, B) (3.25)
m k=0

is an invariant measure for the original chain so that πP= πP . By normalizing this measure, we obtain an invariant
probability measure for the original chain, provided that n nK(n) < ∞ (see Theorem 3.2.3).

Exercise 3.2.1 Show that (3.25) is an invariant probability measure given that (3.24) holds.

3.2.2 Existence of an invariant probability measure

We state the following very consequential results on the existence of invariant probability measures for Markov chains.

Theorem 3.2.5 Consider an aperiodic and irreducible Markov chain {xt }. If there exists a set A which is also an m-small
set for some m ∈ Z+ , and if the set satisfies

sup E[min(t > 0 : xt ∈ A)|x0 = x] < ∞,


x∈A

then the Markov chain admits an invariant probability measure.

Note that for m = 1, we don’t need irreducibility or aperiodicity, by directly following the splitting construction presented.
In the following, we relax aperiodicity, in any case.

Theorem 3.2.6 (Meyn-Tweedie) Consider a Harris recurrent Markov chain {xt }. If there exists a µ-petite set A for some
positive measure µ, and if the set satisfies

sup E[min(t > 0 : xt ∈ A)|x0 = x] < ∞,


x∈A

then the Markov chain is positive Harris recurrent (and admits a unique invariant probability measure).

Remark 3.5. For the m-small case with m > 1, in view of the splitting construction, one question is whether

sup E[min(mt > 0 : xmt ∈ A)|x0 = x] < ∞,


x∈A

or in the petite case with a sampled chain with geometrically sampled times τk , whether
3.2 Uncountable Standard Borel State Spaces 39

sup E[min(τk > 0 : xτk ∈ A)|x0 = x] < ∞,


x∈A

is implied by
sup E[min(t > 0 : xt ∈ A)|x0 = x] < ∞.
x∈A

For the small set case, under irreducibility and aperiodicity, the above holds. See Remark 4.4 on the positive Harris re-
currence discussion for an m-skeleton and split chains: When a Markov chain has an invariant probability measure, the
sampled chain (m-skeleton) also satisfies a drift condition, which then leads to the result that an atom constructed through
an m-skeleton has a finite return property, which can be used to establish the existence of an invariant probability measure.
However, the discussion for the petite set case is more direct and can be arrived at via the properties of a geometrically
P∞along the arguments in Exercise 3.5.4) In particular, an m-small set is 1-small for a geometrically
sampled chain (following
sampled chain (since n=0 P n (x, C)R(n) ≥ R(m)P m (x, C)), which then allows for the arguments for existence to be
applicable more directly (see also [108]). The utilization of petite sets, via a sampled chain, thus allows for relaxing the
aperiodicity requirement and also with a more direct argument as discussed above.

In this case, the invariant measure satisfies the following, which is a generalization of Kac’s Lemma [115]:

Theorem 3.2.7 For a µ-irreducible Markov chain with invariant probability measure π, the following holds:
Z C −1
τX
π(A) = π(dx)Ex [ 1{xk ∈A} ], ∀A ∈ B(X), µ(A) > 0, π(C) > 0
C k=0

Observe that with the above, by taking A = X we have that for any Borel C,
1
inf Ex [τC ] ≤ ≤ sup Ex [τC ]
x∈C π(C) x∈C

3.2.3 On small and petite sets: sufficient conditions (Optional)

Establishing the smallness or petiteness of a set may be difficult to directly verify. In the following, we present a few
conditions that may be used to establish petiteness properties.
T-chains. By [232, p. 131], for a Markov chain with transition kernel P and K a probabilityP∞ measure on natural numbers,
if there exists for every E ∈ B(X), a lower semi-continuous function N (·, E) such that n=0 P n (x, E)K(n) ≥ N (x, E)
for a sub-stochastic kernel N (·, ·) with N (x, X) > 0 for all x ∈ X, the chain is called a T −chain.

Theorem 3.2.8 [232, Theorem 6.2.5] For a T −chain which is irreducible, every compact set S is petite.

Proof Sketch. We will prove the result with thePstronger assumption P (x, A) is continuous in x for every Borel A (that is,

P is strong Feller). Note that this implies that n=0 P n (x, B)K(n) is continuous for every sampling distribution K, by
the dominated convergence theorem. Due to irreducibility, by Theorem 3.2.9, a petite set B exists so that with a positive ν
measure, for every Borel C, we have

X
P n (x, C)R(n) ≥ ν(C), ∀x ∈ B.
n=0
P∞
Now there exists K such that n=0 P n (x, B)K(n) puts a positive measure on B for every x ∈ X, due to irreducibility
and that B has positive measure under the irreducibility measure. Since

X
P n (x, B)K(n)
n=0
40 3 Classification of Markov Chains
P∞
is continuous, there exists x∗ ∈ S such that the minimum n=0 P n (x, B)K(n) over x ∈ S is attained. The desired
petiteness result then comes from bounding, for any Borel C, for an appropriate probability measure η:
X X Z X 
r n ∗ m
η(r)P (x, C) ≥ K(n) P (x , dy) R(m)P (y, C)
r B m

for every x ∈ S. ⋄
A reflection on the proof of the result above, via Lusin’s theorem (see Theorem D.5.1), leads to the following: Small or
petite sets exist for irreducible Markov chains.

Theorem 3.2.9 [232, Thm 5.2.2] Let {xt } be µ-irreducible. Then, for every Borel B with µ(B) > 0, there exists m ≥ 1
and a νm -small set C ⊂ B with µ(C) > 0 and νm (C) > 0.

For a countable state space, under irreducibility, every finite set S is petite.
Tweedie’s uniform countable additivity condition. Tweedie [310] considers the following. If S is such that the following
uniform countable additivity condition

lim sup P (x, Bn ) = 0, (3.26)


n→∞ x∈S

is satisfied for Bn ↓ ∅, then S is petite (and for example, (4.10) to be studied in Chapter 4 implies the existence of an
invariant probability measure). In this case, there exists at most finitely many invariant probability measures. By [232,
Proposition 5.5.5 (iii)], under irreducibility, the Harris recurrent component of the space can be expressed as a countable
union of petite sets Cn with ∪∞ ∞
n=1 Cn , with ∪m Cm → ∅ as m → ∞. By Lemma 4 of Tweedie (2001), under uniform
M
countable additivity, any set ∪i=1 Ci is uniformly accessible from S. Therefore, if the Markov chain is irreducible, the
condition (3.26) implies that the set S is petite. This may be easier to verify for a large class of applications. Under further
conditions (such as if S is compact and V used in a drift criterion (4.10) has compact level sets), then the analysis will lead
sufficient conditions leading to (3.26). In particular, [310, Lemma 1] notes that if S is bounded and V is continuous (and
thus uniformly bounded on S), it suffices to test (3.26) only for Bn sets inside sets on which V is bounded (that is with B1
such that supx∈B1 V (x) < ∞.). In applications, this is often much easier to apply, see e.g. [350].
A further condition. We have the following complementary condition, where irreducibility can be relaxed, but the strong
Feller property is imposed.

Proposition 3.6. [16] Assume that

(i) The transition kernel T is bounded from below by a conditional probability measure that admits a density with respect
to some positive measure ϕ. In other words there exist a measurable f : X × U × X → R+ , such that
Z
P (x, D) ≥ f (x, y)ϕ(dy)
D

for every D ∈ B(X).


(ii) The function f (x, y) is continuous in x for every fixed y.
(iii)It holds that Z
( inf f (x, y))ϕ(dy) > 0
X x∈A
for every nonempty compact set A ⊂ X.

Then, every compact set is 1-small.

Proof. The measurable selection results in Appendix C (see [201, 283] and [168, Theorem 2]) show that, for any compact
A ⊂ X, there exist measurable functions g and F such that
3.2 Uncountable Standard Borel State Spaces 41

inf f (x, y) = min f (x, y) =: F g(y), y (3.27)
x∈A x∈A

Thus, we have for all x ∈ A Z


P (x, D) ≥ inf f (x, y)ϕ(dy)
D x∈A
Z (3.28)
= F (g(y), y)ϕ(dy) =: ν(D)
D
for some finite (sub-probability) measure ν. Thus, every compact set is 1-small. ⋄

3.2.4 Rates of convergence to equilibrium

We can extend Dobrushin’s contraction result for the uncountable state space case. In this general setup, we define the
Dobruhsin coefficient for a Markov chain with transition kernel P as
n
X
δ(P ) = inf min{P (x, Ai ), P (y, Ai )} (3.29)
(x,y);An
i=1

where the infimum is over all x, y ∈ X and all finite partitions An := {Ani , i = 1, · · · , n} consisting of disjoint sets whose
union is X. Note that this definition holds for both continuous or countable X. We then have for two probability measures
π, π ′ (see Appendix D for a review on probability measures) [105]

∥πP − π ′ P ∥T V ≤ (1 − δ(P ))∥π − π ′ ∥T V .

As such, if δ(P ) > 0, the iterations πt = πt−1 P converge to a unique fixed point geometrically fast. To better appreciate
this coefficient, first note that by the property that |a − b| = a + b − 2 min(a, b), the Dobrushin’s coefficient in (3.13) can
be written as (for the countable state space case):
1 X
δ(P ) = 1 − max |(P (i, j) − P (k, j)|
2 i,k j

R the continuous setup, in case P (x, dy) is the transition kernel admitting a density for each x (that is P (x, A) =
For
A
p(x, y)dy with probability density function p(x, ·)), the expression
Z
1
δ(P ) = 1 − sup |p(x, y) − p(z, y)|dy,
2 x,z R

is the Dobrushin’s ergodic coefficient for R−valued Markov processes.


The versatility of using Dobrushin’s coefficient for establishing rates of convergence manifests itself in the following
conditions (noted from [162, Theorem 3.2]).

Theorem 3.2.10 Consider the following conditions.

(i) There exists a state x∗ ∈ X and a number β > 0 such that P ({x∗ }|x) ≥ β for all x ∈ X.
(ii) There exist n ∈ N and a non-trivial (positive )measure µ such that P n (·|x) ≥ µ(·) for all x ∈ X.
(iii)There exist n ∈ N and a positive number β < 1 so that for all x, x′ ∈ X

∥P n (·|x) − P n (·|x′ )∥T V ≤ 2β.

(iv) There exist c > 0, β ∈ (0, 1) such that there is a probability measure π with

∥π0 P n − π∥ ≤ cβ n , π0 ∈ P(X), n ∈ N
42 3 Classification of Markov Chains

We have that
(i) ⇒ (ii) ⇔ (iii) ⇒ (iv)

Note that condition (ii) amounts to the entire state space being n-small. The results above can be established through an
analysis based on Dobrushin’s ergodic coefficient. In the next chapter, we will provide more relaxed conditions leading to
rates of convergence, even though those conditions will not lead to a uniform (over x ∈ X) rate of convergence.

3.3 Further Conditions on the Existence and Uniqueness of Invariant Probability Measures

3.3.1 Further conditions on existence of invariant probability measures

Markov chains with the Feller property

This section uses certain properties of spaces of probability measures, reviewed in Section D.
R
Definition 3.3.1 (i) A Markov chain is weak Feller if X P (x, dz)v(z) is continuous in x for every continuous and bounded
v on X.
R
(ii) If the above holds (i.e., X P (x, dz)v(z) is continuous in x) for every bounded measurable v, the Markov chain is called
strong Feller.

Example 3.7. (i) Let xt+1 = f (xt ) + wt , where f : R → R is continuous and {wt } is an i.i.d. real valued noise sequence.
In this case {xt } is weak Feller, regardless of the random variable wt .
(ii) The chain {xt } is strong Feller if f : R → R is continuous and wt is an i.i.d. random sequence where wt admits a
continuous probability density function.

See Section 5.6 for further examples and discussions (while those examples involve controlled models, you can assume
that the control term is a singleton in the context of the discussion in the current chapter).

Theorem 3.3.1 Let {xt } be a weak Feller Markov process living in a compact subset of a complete, separable metric
space. Then {xt } admits an invariant probability measure.

Proof. Proof follows the observation that the space of probability measures on a compact set is tight (that is, it is weakly
sequentially pre-compact), see Appendix D for a discussion on weak convergence. Consider a sequence
T −1
1 X
µT = µ0 P t , T ≥ 1,
T t=0

There exists a subsequence µTk which converges weakly to some µ∗ . It follows that for every continuous and bounded
function f Z
⟨µTk , f ⟩ := µTk (dx)f (x) → ⟨µ∗ , f ⟩
R
Likewise, since P f (x) = f (x1 )P (dx1 |x0 = x) is continuous in x (by the weak Feller condition), it follows that
Z Z 
⟨µTk , P f ⟩ := µTk (dx) P (x, dy)f (y) → ⟨µ∗ , P f ⟩.

Now,
3.3 Further Conditions on the Existence and Uniqueness of Invariant Probability Measures 43

k −1
 TX k −1
TX 
1
(µTk − µTk P )(f ) = Eµ Pxk f − Pxk+1 f
Tk 0
k=0 k=0
 
1
= Eµ f (x0 ) − f (xTk ) → 0. (3.30)
Tk 0

Thus,
(µTk − µTk P )(f ) = ⟨µTk , f ⟩ − ⟨µTk P, f ⟩ = ⟨µTk , f ⟩ − ⟨µTk , P f ⟩ → ⟨µ∗ − µ∗ P, f ⟩ = 0.
Now, if the relation ⟨µ∗ − µ∗ P, f ⟩ = 0 holds for every continuous and bounded function, it also holds for any measurable
function f : This is because continuous functions are dense in measurable functions under the supremum norm (in other
words, continuous and bounded functions form a separating class for the space of probability measures, see e.g. p. 13
in [42] or Theorem 3.4.5 in [119])). Thus, µ∗ is an invariant probability measure. ⋄

Remark 3.8. The theorem applies identically if instead of a compact set assumption, one assumes that the sequence µk takes
values in a weakly compact set (e.g. via Prohorov’s theorem [110]); that is, if the sequence admits a weakly converging
subsequence.

Remark 3.9. Reference [212] gives the following example to emphasize the importance of the Feller property: Consider a
Markov chain evolving in [0, 1] given by: P (x, x/2) = 1 for all x ̸= 0 and P (0, 1) = 1. This chain does not admit an
invariant measure. This can be established by a continuity of probability argument for any invariant probability measure (if
one existed) on the absorbing sets (0, δ) for any δ > 0.

In the following, we generalize the above result to a case where the state space X is not compact, but is locally compact.

Theorem 3.3.2 Let {xt } be a weak Feller Markov process taking values from a locally compact X. Suppose further for
some initial probability measure µ0 , with
T −1
1 X
µT = µ0 P t , T ≥ 1,
T t=0

we have that for some compact B


lim inf µT (B) > 0.
T →∞

Then, {xt } admits an invariant probability measure.

Proof. The proof builds on an application of the Banach-Alaoglu theorem; the space of signed measures under the total
variation norm with finite total variation is the topological dual of the space of continuous functions which vanish at infinity,
and the unit ball in this space is weak∗ -compact by the Banach-Alaoglu theorem. By an argument similar to the proof of
Theorem 3.3.1, then there exists a subsequence µTk which converges in the weak∗ sense to a limit µ∗ . Since µ∗ cannot
be the trivial (all-zero) measure, µ∗ must be invariant and positive. Normalizing this measure implies that there exists an
invariant probability measure. ⋄

Quasi-Feller chains

Often, one does not have the Feller property, but the set of discontinuity is appropriately negligibe.

Assumption 3.3.1 For f bounded and continuous, P f (x) := E[f (Xt+1 )|Xt = x]) is continuous on X \ D where D is a
closed set with P (Xt+1 ∈ D|x) = 0 for all x. Furthermore, with Dϵ = {z : d(z, D) < ϵ} for ϵ > 0 and d the metric on
X, for some K < ∞, we have that for all x and ϵ > 0
 
P Xt+1 ∈ Dϵ |xt = x ≤ Kϵ.
44 3 Classification of Markov Chains

Theorem 3.3.3 Suppose that Assumption 3.3.1 holds. If the state space is compact, there exists an invariant probability
measure for the Markov chain.

Proof. The sequence of expected empirical probability measures


n−1
1X
vn (A) = Ex [ 1{Xk ∈A} ]
n
k=0

is tight, and thus there exists a weakly converging subsequence. Assumption 3.3.1 implies that every converging subse-
quence vnk of is such that for all ϵ > 0
lim sup vnk (Dϵ ) ≤ Kϵ.
nk →∞

Note that with v = limnk →∞ vnk , it follows from the Portmanteau theorem (see e.g. [110, Thm.11.1.1]) that

v(Dϵ ) ≤ Kϵ.

Now, consider a weakly converging empirical occupation sequence vtk and let this sequence have an accumulation point
v ∗ . We will show that v ∗ is invariant.
Observe that the transitioned probability measure vtk P satisfies the following for every continuous and bounded f : Con-
sider ⟨vtk , P f ⟩ = ⟨vtk , gf ⟩ + ⟨vtk , P f − gf ⟩, where gf is a continuous function which is equal to P f outside an open
neighborhood of D and is continuous with ∥gf ∥∞ = ∥Pf ∥∞ ≤ ∥f ∥∞ . The existence of such a function follows from the
Tietze-Urysohn extension theorem [110], where the closed set is given by X \ Dϵ . It then follows from Assumption 3.3.1
that, for every ϵ > 0 a corresponding gf can be found so that ⟨vtk , P f − gf ⟩ ≤ K∥f ∥∞ ϵ, and since ⟨vtk , gf ⟩ → ⟨v ∗ , gf ⟩,
it follows that

lim sup |⟨vtk , P f ⟩ − ⟨v ∗ , P f ⟩|


tk →∞
= lim sup |⟨vtk , P f − gf ⟩ − ⟨v ∗ , P f − gf ⟩|
tk →∞
≤ lim sup |⟨vtk , P f − gf ⟩| + |⟨v ∗ , P f − gf ⟩|
tk →∞

≤ 2K ′ ϵ (3.31)

Here, K ′ = 2K∥f ∥∞ is fixed and ϵ may be made arbitrarily small. We conclude that v ∗ is invariant. ⋄

Remark 3.10. In his definition for quasi-Feller chains, Lasserre assumes the state space to be locally compact. In the proof
above [344] tightness is invoked directly with no use of convergence properties of the set of functions which decay to zero
as is done in [166]; for a related result see Gersho [141].

Cases without the Feller condition

One can relax the weak Feller condition and instead consider spaces of probability measures which are setwise sequentially
pre-compact. The proof of this result follows from a similar observation as (3.30) but with weak convergence replaced by
setwise convergence (see Appendix D). Note that in this case, if µTk → µ∗ setwise, it follows that µTk P (f ) → µ∗ P (f )
and thus µ∗ is invariant. It can be shown (as in the proof of Theorem 3.3.1) that a (sub)sequence of occupation measures
which converges setwise, converges to an invariant probability measure. A sufficient condition for a sequence of probability
measures to be setwise sequentially compact is that there exists a finite measure π such that vk ≤ π for all k ∈ N [167].
As an example, consider a system of the form:

xt+1 = f (xt ) + wt (3.32)

where wt admits a distribution with a bounded density function, which is positive everywhere and f is bounded. This
system admits an invariant probability measure which is unique.
3.3 Further Conditions on the Existence and Uniqueness of Invariant Probability Measures 45

3.3.2 Uniqueness of an invariant probability measure

Unique ergodicity properties

For a Markov chain, the uniqueness of an invariant probability measure implies the ergodicity of the measure; such a
Markov chain is often referred to as uniquely ergodic.
The following definition will be useful.

Definition 3.11. Let π be a probability measure on X with metric d. The topological support of π is defined with

supp π := {x : π(Br (x)) > 0}, ∀r > 0,

where Br (x) = {y ∈ X : d(x, y) < r}.

Theorem 3.3.4 Let {xt } be a ψ-irreducible Markov chain which admits an invariant probability measure. The invariant
measure is unique.

Proof. Let there be two invariant probability measures µ1 and µ2 . Then, there exists two mutually singular invariant
probability measures ν1 and ν2 , that is ν1 (B1 ) = 1 and ν2 (B2 ) = 1, B1 ∩ B2 = ∅ and that P n (x, B1C ) = 0 for all x ∈ B1
and n ∈ Z+ and likewise P n (z, B1C ) = 0 for all z ∈ B1 and n ∈ Z+ . This then implies that the irreducibility measure has
zero support on B1C and zero support on B2C and thus on X, leading to a contradiction. ⋄
A further result on uniqueness is given next.

Definition 3.12. For a Markov chain with transition kernel P , a point x is accessible (or reachable) if for every y and
every open neighborhood O of x, there exists k > 0 such that P k (y, O) > 0.

One can show that if a point is accessible, it belongs to the (topological) support of every invariant measure (see, e.g.,
Lemma 2.2 in [153]). The support (or spectrum) of a probability measure is defined to be the set of all points x for which
every open neighbourhood of x has positive measure. A Markov chain Vt is said to have the strong Feller property at x if
E[f (Xt+1 )|Xt = x) is continuous at x for every measurable and bounded f .

Theorem 3.3.5 [153] [259] If a Markov chain has the strong Feller property at an accessible point, then the chain can
have at most one invariant probability measure.

Proof. Let there be two invariant probability measures µ1 and µ2 . Then, as earlier, there exists two mutually singular
invariant probability measures ν1 and ν2 , that is ν1 (B1 ) = 1 and ν2 (B2 ) = 1, B1 ∩ B2 = ∅ and that P n (x, B1C ) = 0 for
all x ∈ B1 and n ∈ Z+ and likewise P n (z, B1C ) = 1 for all z ∈ B2 and n ∈ Z+ . Now, every point x in S is so that one
can approach x through two sequences yn , zn , one in B1 and one in B2 whose evaluations of P n (·, B1C ) are 1 apart from
each other as yn , zn converge to one another (through x). This violates strong continuity. ⋄
Another useful result is the following. Let us first recall the following: A family of functions F mapping a metric space S
to R is said to be equi-continuous at a point x0 ∈ S if, for every ϵ > 0, there exists a δ > 0 such that d(x, x0 ) ≤ δ =⇒
|f (x) − f (x0 )| ≤ ϵ for all f ∈ F . The family F is said to be equicontinuous if it is equicontinuous at each x ∈ S.

Definition 3.13. [232, Chapter 6] A Markov chain with transition kernel P is called an e-chain if for each continuous
function f with compact support, the sequence of functions { P n (x, dy)f (y), n ∈ Z+ } is equi-continuous on compact
R

sets.

Theorem 3.3.6 [232, Prop. 18.4.2] If a Markov chain is an e-chain, X is compact, and a reachable state x∗ exists, then
there exists a unique invariant probability measure.
46 3 Classification of Markov Chains

In the following, we present a more concise argument compared with [232, Prop. 18.4.2].
Proof. By compactness, by Theorem 3.3.1 we know that there exists at least one invariant probability measure. Let there
be two different probability measures ν 1 , ν 2 . We may assume ν 1 and ν 2 to be ergodic2 , via an ergodic decomposition
argument of invariant measures on compact subsets [309, Theorem 6.1]. Now, similar to the proof of Theorem 3.3.5, since
x∗ must belong to the support of any two distinct probability measures (recall that Br (x∗ ) is visited under either of the
probability measures in finite time for any given r > 0) we have that there exists two sequences yn , zn which converge to
one another (through x∗ ) where yn , zn belong to the support sets of these two distinct probability measures ν 1 , ν 2 .
Now, by equi-continuity and the Arzela-Ascoli theorem [110], we have that
N −1 Z
1 X
P (N ) (f )(x) := P n (x, dy)f (y) (3.33)
N
k=0

has a subsequence which converges (in the sup norm) to a limit Ff∗ : X → R, where is Ff∗ continuous.
The above imply that, for every continuous and bounded f , the term

lim | lim P (N ) (f )(yn ) − P (N ) (f )(zn )| = 0.


n→∞ N →∞

Suppose not; there would be an ϵ > 0 and a subsequence nk for which the difference

| lim P (N ) (f )(ynk ) − P (N ) (f )(znk )| > ϵ.


N →∞

However, for each fixed nk , we have that


lim P (N ) (f )(ynk )
N →∞

converges by the ergodicity of ν 1 to ⟨ν 1 , f ⟩ and the limit, by the Arzela-Ascoli theorem, will be equal to Ff∗ (ynk ) (as
every converging subsequence would have to converge to the limit; which also implies that the subsequential convergence
in (3.33) is in fact a sequential convergence). The same argument applies for P (N ) (f )(znk ) → ⟨ν 2 , f ⟩ = Ff∗ (znk ).
The above would then imply that |Ff∗ (ynk ) − Ff∗ (znk )| ≥ ϵ for every (ynk , znk ). This would be a contradiction due to the
continuity of Ff∗ .

Therefore, the time averages of f under ν 1 and ν 2 will be arbitrarily close to each other. However, since continuous
functions separate probability measures (e.g. via the metric given in (D.3), see also [119, Theorem 3.4.5]), this implies that
the probability measures ν 1 and ν 2 must be equal. ⋄
There exist further refinements on unique ergodicity via such equi-continuity conditions on transition kernels [194] and
with equi-continuity replaced with uniform Lipschitz regularity [1].

3.4 Ergodic Theorems for Markov Chains

3.4.1 Ergodic theorems for positive Harris recurrent chains

R
Let c ∈ L1 (µ) := {f : |f (x)|µ(dx) < ∞}. Suppose that µ is an invariant ergodic probability measure for a Markov
chain (a sufficient condition being that µ is the unique invariant probability measure [167, Prop. 2.4.3]). Then (see e.g. [167,
Chapter 2]) it follows that for µ almost everywhere x ∈ X:
T Z
1X
lim c(xt ) = c(x)µ(dx),
T →∞ T
t=1

2
we say that an invariant measure µ measure is ergodic if for every absorbing set S, µ(S) ∈ {0, 1} [167, Definition 2.4.1]
3.4 Ergodic Theorems for Markov Chains 47

Px almost surely (that is conditioned on x0 = x, with probability one, the above holds); see also Theorem 3.4.2. Further-
more, again with c ∈ L1 (µ), for µ almost everywhere x ∈ X
XT  Z
1
lim Ex c(xt ) = c(x)µ(dx),
T →∞ T
t=1

On the other hand, the positive Harris recurrence property allows the almost sure convergence to take place for every
initial condition: If µ is the invariant probability measure for a positive Harris recurrent Markov chain, it follows that for
all x ∈ X and for every c ∈ L1 (µ) [232, Theorem 17.1.7] or [167, Theorem 4.2.13]
T Z
1X
lim c(xt ) = c(x)µ(dx), (3.34)
T →∞ T
t=1

almost surely. However, for every c ∈ L1 (µ), while (7.48) holds for all x ∈ X, it is not generally true that
T Z
1 X
lim Ex [ c(xt )] = c(x)µ(dx).
T →∞ T
t=1

Thus, we can not in general relax the boundedness condition for the convergence of the expected costs. However, with c
bounded, forall x ∈ X
XT  Z
1
lim Ex c(xt ) = c(x)µ(dx) (3.35)
T →∞ T
t=1

This follows as a consequence of Fatou’s lemma and (7.48). Further refinements are possible via return properties to small
sets and f -regularity of cost functions [13, 232]; e.g., this convergence holds if Theorem 4.2.4 holds for the given f and
for some Lyapunov function V with X0 = x ∈ {z : V (z) < ∞}; see the proof of Theorem 4.2.4 and [232, Theorem
14.0.1] for further related results. We refer the reader to [232, Chapters 14 and 17] or [167, Chapters 2 and 4] for additional
discussions. See Exercise 3.5.8.

3.4.2 Further ergodic theorems for Markov chains

Although beyond the scope of this course, for completeness, we state the following. When an invariant probability measure
is known to exist for a Markov chain, we state the following ergodicity results.

Theorem 3.4.1 [166, Theorems 2.3.4-2.3.5] Let P̄ be an invariant probability measure for a Markov process.

(i) [Individual ergodic theorem] Let X0 = x. For every f ∈ L1 (P̄ )


N −1
1 X
Ex [ f (Xn )] → f ∗ (x),
N n=0

for all x ∈ Bf where P̄ (Bf ) = 1 (where Bf denotes that the set of convergence may depend on f ) for some f ∗ .
PN −1
(ii) [Mean ergodic theorem] Furthermore, the convergence N1 Ex [ n=0 f (Xn )] → f ∗ (x) is in L1 (P̄ ).

Theorem 3.4.2 [166, Theorem 2.5.1] Let P̄ be an invariant probability measure for a Markov process. With X0 = x, for
every f ∈ L1 (P̄ )
N −1
1 X
f (Xn ) → f ∗ (x),
N n=0
48 3 Classification of Markov Chains

for all x ∈ Bf where P̄ (Bf ) = 1 for some f ∗ (x) with


Z Z

P̄ (dx)f (x) = P̄ (dx)f (x)

One may state further refinements; see [166] for the locally compact case and [336] for the Polish state space case.

Theorem 3.4.3 [336, Prop. 5.4] or [167, Theorem 3.1(g)] Let P̄ be an invariant probability measure for a Markov
process.
1
PN −1
(i) [Ergodic decomposition and weak convergence] For x, P̄ a.s., N Ex [ t=0 1{xn ∈·} ] → Px (·) weakly and P̄ is invari-
ant for Px (·) in the sense that Z
P̄ (B) = Px (B)P̄ (dx)

(ii) [Convergence in total variation] For all µ ∈ P(X) which satisfies that µ ≪ P̄ (that is, µ is absolutely continuous with
respect to P̄ ), there exists an invariant v ∗ such that
N −1
1 X
∥Eµ [ 1{T n X∈·} ] − v ∗ (·)∥T V → 0.
N t=0

3.5 Exercises

Exercise 3.5.1 For a countable state space Markov chain, prove that if {xt } is irreducible, then all states have the same
period.

Exercise 3.5.2 Prove that


Px (τA = 1) = P (x, A),
and for n ≥ 1, X
Px (τA = n) = P (x, i)Pi (τA = n − 1)
i∈A
/

Exercise 3.5.3 Let {xt } be a Markov chain defined on state space {0, 1, 2}. Let the one-stage probability transition matrix
be given by:  
0 1 0
P = 1/2 0 1/2
0 2/3 1/3

Compute E[min(t ≥ 0 : xt = 2)|x0 = 0], that is the expected minimum number of stages for the state to move from 0 to 2.
Hint: Building on the previous exercise, one way to solve this problem is as follows: Note that if the expected minimum
time to go to state 2 is from state 1 is t1 and the expected minimum time to go to state 2, from state 0 is t0 , then the expected
minimum time to go to state 2 from state 0 will be t0 = 1 + P (0, 2)t2 + P (0, 1)t1 + P (0, 0)t0 , where t2 = 0. You can
follow this line of reasoning to obtain the result.

Exercise 3.5.4 Let (Ω, F, P ) be a probability space on which a Markov chain is defined: Let X be a finite set and Xn be
the X-valued Markov chain. Let α ∈ X with Eα [τα ] = 5, where

τα := min{k > 0 : xk = α}

As we know from our class, this Markov chain admits an invariant probability measure, call it π. Suppose that Xn is
irreducible so that the invariant probability measure is unique.
3.5 Exercises 49

Now, let Yn be an i.i.d. {0, 1}-valued Bernoulli process (defined on the same probability space) with

P (Yn = 1) = η ∈ (0, 1).

Let
τ(α,1) := min{k > 0 : (Xk , Yk ) = (α, 1)}.
a) Is Zk := (Xk , Yk ) a Markov chain? Prove your answer.
b) Find E[τ(α,1) |(X0 , Y0 ) = (α, 1)].

Exercise 3.5.5 Show that irreducibility of a Markov chain in a finite state space implies that every set A and every x
satisfies U (x, A) = ∞.

Exercise 3.5.6 Show that for an irreducible Markov chain, either the entire chain is transient, or recurrent.

Exercise 3.5.7 Show that for αt ∈ (0, 1),



Y
(1 − αt ) > 0
t=0
P
if and only if t αt < ∞.

Hint. For one direction, use log(1 − x) < −x for small x ∈ (0, 1). ForPthe other direction, use limx→0 log(1−x)
x = −1 and
that as a result for small enough x > 0 : log(1 − x) > −2x and that t αt < ∞ implies that αt → 0.

Exercise 3.5.8 In view of Exercise 3.5.7, let us revise the example given in Remark 3.2: Let X = N, P (1, 1) = 1 and for
x > 1: P (x, x + 1) = 1 − 1/x and P (x, 1) = 1/x. This chain is then Positive Harris Recurrent with invariant measure δ1
and irreducibility measure also δ1 . Show that with f (x) = x − 1:
N −1
1 X
lim Ex [ f (xk )] ̸= Eδ1 [f (X)] = 0, x ̸= 1
N →∞ N
k=0

Thus, expected empirical summations do not necessarily converge to the summation under the invariant measure when
the function is not bounded. Note that this would be the case if the functions are bounded. Observe also that sample path
convergence here does not imply the convergence of expected averages.

P∞ 3.5.9 For a Markov chain with a countable space X, and a ∈ X, show that if Pa (τa < ∞) < 1 then
Exercise
Ea [ k=1 1{xk =a} ] < ∞.

P∞ 3.5.10 For a Markov chain with a countable space X, and a ∈ X, show that if Pa (τa < ∞) = 1 then
Exercise
Pa ( k=1 1{xk =a} = ∞) = 1.

Exercise 3.5.11 Consider a Markov chain with state space [0, 1] and transition kernel given as follows:
x
P (X1 = |X0 = x) = 1, x ∈ [0, 1].
4
Does there exist an invariant probability measure π for this Markov chain? If so, what is one such measure? Is this a unique
invariant probability measure? If there is no invariant probability measure, precisely explain why this is the case.

Exercise 3.5.12 (Gambler’s Ruin) Consider an asymmetric random walk defined as follows: P (xt+1 = x + 1|xt = x) =
p and P (xt+1 = x − 1|xt = x) = 1 − p for any integer x. Suppose that x0 = x is an integer between 0 and N . Let
τ = min(k > 0 : xk ∈
/ [1, N − 1]). Compute Px (xτ = N ) (you may use Matlab for your solution).
Hint: Observe that one can obtain a recursion as Px (xτ = N ) = pPx+1 (xτ = N ) + (1 − p)Px−1 (xτ = N ) for
1 ≤ x ≤ N − 1 with boundary value conditions PN (xτ = N ) = 1 and P0 (xτ = N ) = 0. One observes that
50 3 Classification of Markov Chains
 
1−p
Px+1 (xτ = N ) − Px (xτ = N ) = Px (xτ = N ) − Px−1 (xτ = N )
p

and in particular
 
1 − p N −1
PN (xτ = N ) − PN −1 (xτ = N ) = ( ) P1 (xτ = N ) − P0 (xτ = N )
p

Exercise 3.5.13 Let xt+1 = f (xt ) + wt , where f : R → R is continuous and {wt } is an i.i.d. real valued noise sequence.
a) Show that {xt } is weak Feller, regardless of the random variable wt .
b) Show that {xt } is strong Feller, if wt is a Gaussian random variable with a positive variance.

Exercise 3.5.14 a) Consider a Markov chain defined on Z+ with the transition kernel

1
P (x1 = x + 1|x0 = x) = 1 − , x ̸= 0, x ∈ Z+ ,
x+1
1
P (x1 = 0|x0 = x) = , x ̸= 0, x ∈ Z+ ,
x+1
with
P (x1 = 1|x0 = 0) = 1.

Does there exist an invariant probability measure π for this Markov chain? If so, what is one such measure?
b) Consider a Markov chain defined on Z+ with the transition kernel

1
P (x1 = x + 1|x0 = x) = , x ̸= 0, x ∈ Z+ ,
x+1
1
P (x1 = 0|x0 = x) = 1 − , x ̸= 0, x ∈ Z+ ,
x+1
with
P (x1 = 1|x0 = 0) = 1.

Does there exist an invariant probability measure π for this Markov chain? If so, what is one such measure?
b) Consider a Markov chain defined on [0, 1] with the transition kernel:
x
P (x1 = |x0 = x) = 1, x ̸= 0, x ∈ [0, 1],
2
P (x1 = 1|x0 = 0) = 1.

Does there exist an invariant probability measure π for this Markov chain? If so, what is one such measure?

Exercise 3.5.15 Consider a square and join opposite corners of this square by straight lines meeting at the point C.
Consider the symmetric random walk performed by a particle on these 5 vertices, starting at some vertex A. Find
(a) the expected time to return to A,
(b) the expected number of visits to C before returning to A,
(c) the expected time to return to A given that there is no prior visit to C.
4

Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and


Stochastic Iterative Dynamics

In this chapter, we will study martingales, which constitute a critical class of stochastic processes for our understanding
of stochastic dynamics. We will arrive at stochastic stability of Markov chains through martingale methods and Foster-
Lyapunov type stability criteria; this will be followed by an analysis on stochastic iterative dynamics.

4.1 Martingales

In this section, we introduce martingales and discuss a number of important martingale theorems. Only a few of these will
be critical within the scope of our coverage, some others are presented for completeness.
These are very important for us to understand stabilization of controlled stochastic systems. These also will pave the way
to optimization of dynamical systems as well as the supporting theory for stochastic learning, reinforcement learning, and
approximation algorithms to be studied later. The second half of this chapter is on the stability of Markov chains or the
stabilization of controlled Markov Chains via martingale and Lyapunov methods.

4.1.1 More on expectations and conditional probability

Let (Ω, F, P ) be a probability space and let G be a subset of F which is itself a σ-field (such a collection is said to be
a sub-σ-field of F). Let X be an R−valued random variable measurable with respect to (Ω, F) with a finite absolute
expectation that is Z
E[|X|] = |X(ω)|P (dω) < ∞,

where ω ∈ Ω. We call such random variables integrable.
We say that Ξ is the conditional expectation random variable (and is also called a version of the conditional expectation)
of X given G, denoted with,
E[X|G],
if

1. Ξ is G-measurable.
2. For every A ∈ G,
E[1A Ξ] = E[1A X],
where we use the notation Z Z
E[1A Ξ] = Ξ(ω)1{ω∈A} P (dω) = Ξ(ω)P (dω)
Ω A
52 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

For example, if the information that we know about a process is whether an event A ∈ F happened or not -that is, the
sub-σ-field is σ({A}) = {∅, Ω, A, Ω \ A}-, then the conditional expectation of X given the sigma-field generated by A is
the random variable:
η(ω) = E[X|σ({A})](ω) = XA 1{ω∈A} + XAC 1{ω∈A} / .
Observe that this random variable is σ({A})-measurable. We then have that for w ∈ A:
Z
1
η(ω) =: XA =: E[X|A] = P (dω)X(ω).
P (A) A
R R
This follows from the fact that A E[X|A](ω)P (dω) = E[X|A] A P (dω) since η(ω) = E[X|A] cannot distinguish
between any ω ∈ A: this is a consequence of the fact that the conditional expectation is σ({A})-measurable. Thus, we can
simply write η(ω) = E[X|A] for the conditional expectation rather that η(ω). If the information we have is that A did not
take place; then for ω ∈
/ A:
Z
C 1
η(ω) =: XAC := E[X|A ] = P (dω)X(ω).
P (Ω \ A) Ω\A

Note that conditional probability can be expressed as

P (X ∈ B|G) = E[1{X∈B} |G],

hence, conditional probability is a special case of conditional expectation.


It is a useful exercise (see Exercise 1.6.6) to consider the σ-field generated by an observation variable, and what a condi-
tional expectation means in this case.

Theorem 4.1.1 Let X be an X valued random variable, where X is a complete, separable, metric space and Y be another
Y−valued random variable, Then, X is FY or σ(Y ) (the σ−field generated by Y ) measurable if and only if there exists a
measurable function f : Y → X such that X(ω) = f (Y (ω)).

With the above, the expectation E[X|σ(Y )](ω : Y (ω) = y0 ) = E[X|Y = y0 ] can be defined as a measurable function on
σ(Y ), and this expectation can be expressed as a measurable function of Y .
The notion of conditional expectation is key for the development of stochastic processes which evolve according to a
transition kernel. This is useful for optimal decision making when only partial information is available with regard to a
random variable.
The following discussion is optional until the next subsection.

Existence of Conditional Expectation.

Theorem 4.1.2 (Radon-Nikodym) Let µ and ν be two σ-finite positive measures on (Ω, F) such that ν(A) = 0 implies
that µ(A) = 0 (that is µ is absolutely continuous with respect to ν). Then, there exists a measurable function f : Ω → R+ ,
(integrable under ν if µ is a finite measure), such that for every A:
Z
µ(A) = f (ω)ν(dω)
A

The representation above is unique, up to sets of measure zero. With the above discussion, the conditional expectation
X = E[X|F ′ ] exists for any sub-σ-field F ′ ⊂ F, as the following discussion shows. Let X be an integrable non-negative
random variable and observe that for any Borel A ∈ F ′
Z   Z
E[X|F ′ ](ω) P (dω) = X(ω)P (dω).
A A
4.1 Martingales 53

X(ω)P (dω)as a measure(defined on the measurable space (Ω, F ′ )) which is absolutely


R
We may view ζ(A) := A

continuous with respect to P , and thus, E[X|F ′ ](ω) , is the Radon-Nikodym derivative of this measure with respect
to P (This discussion extends to arbitrary integrable variables by considering the negative valued portion of the variable
separately).
In case X is a real random variable which is of second-order (i.e., with finite second moment), another way to establish
existence is through a Hilbert theoretic approach, by viewing the conditional expectation as the projection of X onto a
subspace consisting of the set of all functions measurable on F ′ . We will revisit this later in the notes while deriving
the Kalman Filter in Chapter 6. However, for this we would require X to be square-integrable (i.e., with a finite second
moment).

4.1.2 Some properties of conditional expectation:

One very important property is given by the following.


Iterated expectations:

Theorem 4.1.3 For three σ-fields over a given set, if H ⊂ G ⊂ F, and X is F−measurable and integrable, it follows
that:
E[E[X|G]|H] = E[X|H]

Proof. Proof follows by taking a set A ∈ H, which is also in G and F. Let η be the conditional expectation variable with
respect to H. Then it follows that
E[1A η] = E[1A X]
Now let E[X|G] be η ′ . Then, it must be that E[1A η ′ ] = E[1A X] for all A ∈ G and hence for all A ∈ H. Thus, we have
that for all A ∈ H
E[1A η ′ ] = E[1A η],
and as η is H-measurable, E[η ′ |H] = η. ⋄

Theorem 4.1.4 Let G ⊂ F, and Y be G−measurable. Let X be F−measurable and XY be integrable. Then, P almost
surely
E[XY |G] = Y E[X|G]
Pn
Proof. First assume that Y = Yn is a simple function (a simple random variable of the form: Yn (ω) = i=1 ai 1{ω∈Ai }
with Ai ∈ G). Let us call E[X|G] = η and call E[XY |G] = ζ.
Then, for all A ∈ G
Z n
Z X
Yn η(ω)P (dω) = ai 1{ω∈Ai } η(ω)P (dω)
A A i=1
n
X Z n
X Z
= ai η(ω)P (dω) = ai X(ω)P (dω) (4.1)
i=1 A∩Ai i=1 A∩Ai
n
Z X Z
= ai 1{ω∈Ai } X(ω)P (dω) = Yn XP (dω)
A i=1 A

Here, (4.1) holds since A ∩ Ai ∈ G and E[X|G] = η. On the other hand,


Z Z
ζ(ω)P (dω) = X(ω)Yn (ω)P (dω)
A A
54 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics
n
Z X n
X Z Z
= ai 1{ω∈Ai } X(ω)P (dω) = ai X(ω)P (dω) = Yn XP (dω)
A i=1 i=1 A∩Ai A

Thus, for all A ∈ G


Z Z Z
Yn XP (dω) = E[XYn |G](ω)P (dω) = Yn E[X|G](ω)P (dω), (4.2)
A A A

and the two conditional expectations E[XYn |G] and Yn E[X|G] are equal. Now, the proof is complete by noting that
any integrable Y can be approached from below monotonically by a sequence of simple functions measurable on G. The
monotone convergence theorem leads to the desired result. ⋄

4.1.3 Discrete-time martingales

Let (Ω, F, P ) be a probability space. An increasing family {Fn } of sub-σ−fields of F is called a filtration.
A sequence of random variables defined on (Ω, F, P ) is said to be adapted to Fn if Xn is Fn -measurable, that is
Xn−1 (D) = {w ∈ Ω : Xn (w) ∈ D} ∈ Fn for all Borel D. This holds for example if Fn = σ(Xm , m ≤ n), n ≥ 0; in this
case we call the filtration, the natural filtration.
Given a filtration Fn and a sequence of real random variables adapted to it, (Xn , Fn ) is said to be a martingale if

E[|Xn |] < ∞

and
E[Xn+1 |Fn ] = Xn .
We will often take the sigma-fields to be the natural filtration Fn = σ(X1 , X2 , . . . , Xn ).
Let n > m ∈ Z+ . Since Fm ⊂ Fn , it must be that A ∈ Fm should also be in Fn . Thus, if Xn is a martingale sequence,
we have for A ∈ Fm
E[1A Xn ] = E[1A Xn−1 ] = · · · = E[1A Xm ],
and thus E[Xn |Fm ] = Xm .
If we have that
E[Xn |Fm ] ≥ Xm
then {Xn } is called a submartingale.
And, if
E[Xn |Fm ] ≤ Xm
then {Xn } is called a supermartingale.
A useful concept related to filtrations is that of a stopping time, which we discussed while studying Markov chains. A
stopping time τ is a random time, whose occurrence at any given time is causally measurable with respect to the filtration
in the sense that for each n ∈ N, {τ ≤ n} ∈ Fn .

Definition 4.1.1 (Filtration up to a stopping time) Let Ft denote a filtration and τ be a stopping time with respect to this
filtration so that for every k, {τ ≤ k} ∈ Fk . Then, the σ-field of events up to τ , Fτ , is the collection of all events A ∈ F
that satisfies:
A ∩ {τ ≤ t} ∈ Ft , ∀t ∈ Z+ .

Intuitively, then, the natural filtration up to a stopping time is all the information generated by a stochastic process up to
the stopping time.

Exercise 4.1.1 Let τ be a stopping time with τ ≥ m for some m ∈ Z+ . Show that Fm ⊂ Fτ .
4.1 Martingales 55

4.1.4 Doob’s optional sampling theorem

Theorem 4.1.5 Suppose (Xn , Fn ) is a martingale sequence, and ρ, τ < n (for some fixed n ∈ N) are (uniformly) bounded
stopping times with ρ ≤ τ . Then,
E[Xτ |Fρ ] = Xρ

Proof. We observe that


τ
X −1
E[Xτ − Xρ |Fρ ] = E[ Xk+1 − Xk |Fρ ]
k=ρ

X
= E[ 1{τ >k} (Xk+1 − Xk )|Fρ ]
k=ρ
Xn
= E[ 1{τ >k} (Xk+1 − Xk )|Fρ ]
k=ρ
n
X
= E[1{τ >k} (Xk+1 − Xk )|Fρ ] (4.3)
k=ρ
Xn
= E[E[1{τ >k} (Xk+1 − Xk )|Fk ]|Fρ ]
k=ρ
n
X
= E[ 1{τ >k} E[(Xk+1 − Xk )|Fk ]|Fρ ]
k=ρ
τ
X −1
= E[ 0|Fρ ] = 0 (4.4)
k=ρ

Here, we invoke Theorem 4.1.3 and Theorem 4.1.4, since 1{τ >k} is Fk -measurable. ⋄
The statement of the theorem leads to inequalities for supermartingales or submartingales with the appropriate inequality
signs.
In the above, the main properties we used were (i) the fact that the sub-fields are nested, (ii) n is bounded so that the
expectation of the sum can be written as the sum of expectations in (4.3).
Let us try to see why boundedness of the stopping times is important: Consider the following game. Suppose that one
draws a fair coin; with equal probabilities of heads and tails. If we have a tail, we win a dollar, and a head will cause
us to lose a dollar. Suppose we have 0 dollars at time 0 and we decide to stop when we have 5 dollars, that is at time
τ = min(n > 0 : Xn = 5). In this case, clearly E[Xτ ] = 5, as we will stop when we have 5 dollars. But E[Xτ ] ̸= X0 !

Remark 4.1. For this example, Xn = Xn−1 + Wn , where Wn is either −1 or 1 with equal probabilities and Xn is the
amount of money we have. Clearly Xn is a martingale sequence. The problem is that one might have to wait for an
Pn long period of amount of time to be able to have the 5 dollars, the sequence E[|Xn |] is not uniformly bounded,
arbitrarily
and k=ρ 1{τ >k} (Xk+1 − Xk ) = Xmin(τ,n) − Xρ is not a (uniformly) integrable sequence, and the proof method adopted
in Theorem 4.1.5 will not be applicable. Note that if we were able to claim that

E[Xτ − Xρ |Fρ ] = E[ lim Xmin(τ,n) − Xρ |Fρ ]


n→∞
 X n   n
X
= E lim 1{τ >k} (Xk+1 − Xk ) Fρ = lim E[ 1{τ >k} (Xk+1 − Xk )|Fρ ], (4.5)
n→∞ n→∞
k=ρ k=ρ

then the result would have been applicable even if we didn’t have a finite upper bound on the stopping times. This requires
in particular:
56 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

(i) the almost sure finiteness of τ so that limn→∞ Xmin(τ,n) = Xτ , and


Pn
(ii) a dominated convergence result for the sequence Xmin(τ,n) − Xρ = k=ρ 1{τ >k} (Xk+1 − Xk ) (e.g., the presence of
an integrable random variable G(ω) with Xmin(τ,n) (ω) ≤ G(ω)).

If these hold, then we can indeed apply the argument above when the stopping times are not bounded. More on this will be
discussed below in Theorem 4.1.14, in the context of uniform integrability.

4.1.5 Doob’s maximal inequality (optional)

Theorem 4.1.6 For a non-negative supermartingale Mn , for all λ > 0,


M0
P ( sup Mn ≥ λ) ≤
0≤n<∞ λ

 
N
Proof. Let τ = min min{n ≥ 0 : Mn ≥ λ}, N for some N ∈ N. Then,

E[Mτ N ] M0
P ( max Mn ≥ λ) = P (Mτ N ≥ λ) ≤ ≤ ,
0≤n<N λ λ
where the first inequality follows from Markov’s inequality and the last from Doob’s optional sampling theorem. The
relation above applies for all N ∈ N, and since the left hand side is non-decreasing in N , the limit of it as N → ∞
is well-defined. Furthermore, by an application of continuity in probability limN →∞ P (max0≤n<N Mn ≥ λ) =
P (sup0≤n<∞ Mn ≥ λ), and the result follows. ⋄
An important generalization is known as Doob’s Lp -maximal inequality.


Theorem 4.1.7 For a submartingale Mn , let MN = sup0≤n<N |Mn |. Then, for every p > 1:

∗ p p p
E[|MN | ]≤( ) E[|MN |p ]
p−1

∗ p
R∞ ∗ p
R∞ ∗ p
Proof Sketch. Write E[|MN | = 0 P (|MN | > t)dt = p 0 P (|MN | > sp )sp−1 ds, where we apply a change of
p N
variables with t = s . Let τs := min{min{n ≥ 0 : Mn ≥ s}, N }. Then, we have that for every s > 0,

sP (|MN | > s) ≤ E[|MτsN |1{|MN∗ |>s} ] ≤ E[|MN |1{|MN∗ |>s} ],

by the submartingale property and the optional sampling theorem (as E[|MN ||FτsN ] ≥ |E[MN |FτsN ]| ≥ |MτsN |). There-
fore,
Z ∞ Z ∞
∗ p ∗
E[|MN | ≤p P (|MN | > s)sp−1 ds ≤ p E[|MN |1|MN∗ |>s ]sp−2 ds
0 0

Z ∞ Z |MN |
= E[|MN | p1{|MN∗ |>s} sp−2 ds] = E[|MN | psp−2 ds]
0 0
p p 1 1
∗ p 1− p
= E[|MN | |M ∗ |p−1 ] ≤ (E[|MN |p ]) p (E[|MN | ]) , (4.6)
p−1 N p−1
where we use Hölder’s inequality in the final step. The result then follows by rearranging the terms. ⋄
This inequality is very useful in the approximation theory for controlled diffusions, as it relates the maximal deviations to
Lp deviations.
4.1 Martingales 57

4.1.6 An important martingale convergence theorem

We first discuss Doob’s upcrossing lemma. Let (a, b) be a non-empty interval. Let X0 ∈ (a, b). Define a sequence of
stopping times

T1 = min{N ; min(n : 0 ≤ n ≤ N, Xn ≤ a)} T2 = min{N ; min(n : T1 ≤ n ≤ N, Xn ≥ b)}

T3 = min{N ; min(n : T2 ≤ n ≤ N, Xn ≤ a)} T4 = min{N ; min(n : T3 ≤ n ≤ N, Xn ≥ b)}


and for m > 2:

T2m−1 = min{N ; min(n : T2m−2 ≤ n ≤ N, Xn ≤ a)} T2m = min{N ; min(n : T2m−1 ≤ n ≤ N, Xn ≥ b)}

The number of upcrossings of (a, b) up to time N is the random variable ζN (a, b) = the number of times between 0 and
N , {Xn } crosses the strip (a, b) from below a to above b.
Note that if the sequence is a supermartingale, XT2 − XT1 has a negative expectation.

Theorem 4.1.8 Let XT be a supermartingale sequence. Then,

E[max(0, a − XN )] E[|XN |] + |a|


E[ζN (a, b)] ≤ ≤ .
b−a b−a

Proof.
There are three possibilities that might take place: The process can end, at time N while the process is below a, between a
and b, or above b. If it crosses above b, then we have completed an upcrossing. In view of this, we may proceed as follows:
Let βN := min(m : T2m = N or T2m−1 = N ) (note that if T2m−1 = N , T2m = N as well). By the supermartingale
property

XβN
0 ≥ E[ XT2i − XT2i−1 ]
i=1
βN
X  βN
X 
= E[ XT2i − XT2i−1 1{T2βN −1 ̸=N } 1{T2βN =N } ] + E[ XT2i − XT2i−1 1{T2βN −1 =N } 1{T2βN =N } ]
i=1 i=1
N −1
 βX 
=E XT2i − XT2i−1 + E[(XN − XT2βN −1 )1{T2βN −1 ̸=N } 1{T2βN =N } ] (4.7)
i=1

Thus,
N −1
 βX 
E XT2i − XT2i−1 ≤ −E[(XN − XT2βN −1 )1{T2βN −1 ̸=N } 1{T2βN =N } ]
i=1
= E[(XT2βN −1 − XN )1{T2βN −1 ̸=N } 1{T2βN =N } ] ≤ E[max(0, a − XN )1{T2βN =N } ] ≤ E[max(0, a − XN )] (4.8)
 
PβN −1
Since, E i=1 XT2i − XT2i−1 ≥ E[βN − 1](b − a), it follows that ζN (a, b) = (βN − 1) satisfies:

E[ζN (a, b)](b − a) ≤ E[max(0, a − XN )] ≤ |a| + E[|XN |],

and the result follows. ⋄


Recall that a sequence of random variables Xn defined on a probability space (Ω, F, P ) converges to X almost surely (a.
s.) if  
P w : lim Xn (w) = X(w) = 1.
n→∞
58 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

Theorem 4.1.9 Suppose Xn is a supermartingale and supn≥0 E[max(0, −Xn )] < ∞. Then limn→∞ Xn = X exists
(almost surely). The same result applies for submartingales, by regarding −Xn as a supermartingale and the condition
supn≥0 E[max(0, Xn )] < ∞. A sufficient condition for both cases is that

sup E[|Xn |] < ∞.


n≥0

Proof. The proof follows from Doob’s upcrossing lemma. Now, for any fixed a, b (independent of ω) with a < b, by the
upcrossing lemma we have that

E[|XN |] + |a|
E[ζN (a, b)] ≤ E[max(0, a − XN )] ≤ ,
(b − a)

which is uniformly bounded. The above holds for every N . Since ζN (a, b) is a monotonically increasing sequence in N ,
by the monotone convergence theorem it follows that

E[|XN |] + |a|
lim E[ζN (a, b)] = E[ lim ζN (a, b)] ≤ sup < ∞.
N →∞ N →∞ N (b − a)

Thus, for every fixed a < b, the number of up-crossings has to be finite almost surely. Hence, the limsup cannot be above
b and the liminf cannot be below a, for otherwise the number of up-crossings would be infinite. It then follows that

P (ω : | lim sup Xn (ω) − lim inf Xn (ω)| > (b − a)) = 0,

since this probability can be expressed also as


  
P ∪r∈Q ω : lim sup Xn (ω) > (b − a + r), lim inf Xn (ω)| < r ,

and by a union bound argument, the probability is upper bounded by the probability of a countable union of zero probability
events which is zero. Finally, a continuity of probability argument (for taking b − a → 0) then leads to
 
P ω : | lim sup Xn (ω) − lim inf Xn (ω)| > 0 = 0.


We can also show that the limit variable has finite absolute expectation.

Theorem 4.1.10 (Submartingale Convergence Theorem) Suppose Xn is a submartingale and supn≥0 E[|Xn |] < ∞.
Then X := limn→∞ Xn exists (almost surely) and E[|X|] < ∞.

Proof. Note that, supn≥0 E[|Xn |] < ∞, is a sufficient condition both for a submartingale and a supermartingale in Theo-
rem 4.1.9. Hence Xn → X almost surely. For finiteness, suppose E[|X|] = ∞. By Fatou’s lemma,

lim sup E[|Xn |] ≥ lim inf E[|Xn |] ≥ E[lim inf |Xn |] = E[ lim |Xn |] = ∞.
n→∞ n→∞ n→∞ n→∞

But this is a contradiction as we had assumed that supn E[|Xn |] < ∞. ⋄

4.1.7 The ergodic theorem[Optional]

See Exercise 4.5.11.


4.1 Martingales 59

4.1.8 Further martingale theorems [Optional]

This section is optional. If you wish not to study it, please proceed to the discussion on stabilization of Markov Chains.

Theorem 4.1.11 Let Xn be a martingale such that Xn converges to some integrable X in L1 that is E[|Xn − X|] → 0.
Then,
Xn = E[X|Fn ], n ∈ N

We will use the following while studying the convex analytic method, as well as on the stabilization of Markov chains while
extending the optional sampling theorem to situations where the sampling (stopping) time is not bounded from above by a
finite number. Let us define uniform integrability:

Definition 4.1.2 : A sequence of random variables {Xn , n ∈ N} is uniformly integrable if


Z
lim sup |Xn |P (dXn ) = 0
K→∞ n∈N |Xn |≥K

This implies that


sup E[|Xn |] < ∞
n

Let for some ϵ > 0,


sup E[|Xn |1+ϵ ] < ∞.
n

This implies that the sequence is uniformly integrable as

|Xn | ϵ
Z Z
1
sup |Xn |P (dXn ) ≤ sup ( ) |Xn |P (dXn ) ≤ sup ϵ E[|Xn |1+ϵ ] → 0,
n |Xn |≥K n |Xn |≥K K n K

as K → ∞. The following result is important in many applications:

Theorem 4.1.12 If Xn is a uniformly integrable martingale, then X = limn→∞ Xn exists almost surely (for all sequences
with probability 1) and in L1 (i.e. E[|X − Xn |] → 0), and Xn = E[X|Fn ].

Theorem 4.1.13 Let X be integrable and Fn be a filtration (not necessarily the natural filtration). Then, Mn = E[X|Fn ]
is uniformly integrable.

Proof. First note that by Jensen’s inequality |E[X|Fn ]| ≤ E[|X||Fn ] (since | · | is a convex function). Therefore, by
Markov’s inequality, for any K ∈ R+ :

E[|X|]
P (|E[X|Fn ]| > K] ≤ E[|E[X|Fn ]|]/K ≤ E[E[|X||Fn ]]/K = ,
K
which decays to zero as K → ∞. Now, consider the measure defined with |X(ω)|dP (ω): For any set sequence Am with
P (Am ) → 0, we have that

lim E[1Am |X|] = 0 (4.9)


m→∞

This follows from a contradiction argument: suppose this is not true,Pthen there exists a subsequence Amk with E[1Amk |X|] ≥
ϵ some fixed ϵ > 0 and a further subsequence Am′k with a finite m′ P (Am′k ). Then a monotone convergence theorem
k
violation can be established so that with Bn = ∪m′k ≥n Am′k , E[1Bn |X|] ̸→ 0 where Bn is a monotone decreasing sequence
whose measure vanishes. Therefore
   
E |E[X|Fn ]|1{|E[X|Fn ]|>K} ≤ E E[|X||Fn ]1{|E[X|Fn ]|>K}
60 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics
 
= E E[|X|1{|E[X|Fn ]|>K} |Fn ] = E[|X|1{|E[X|Fn ]|>K} ]

where we use Theorem 4.1.3 (iterated expectations) and thus combining the above
   
lim sup E |E[X|Fn ]|1|E[X|Fn ]|>K ≤ lim sup E |X|1|E[X|Fn ]|>K = 0,
K→∞ n K→∞ n

where the last step follows from (4.9). ⋄

Optional Sampling Theorem For Uniformly Integrable Martingales

The following builds on Remark 4.1.

Theorem 4.1.14 Let (Xn , Fn ) be a uniformly integrable martingale sequence, and ρ, τ are finite stopping times with
ρ ≤ τ . Then,
E[Xτ |Fρ ] = Xρ

Proof. See the discussion following (4.5) for an explicit analysis and derivation. ⋄

Azuma-Hoeffding inequality for martingales with bounded increments

The following is an important concentration result:

Theorem 4.1.15 Let Xt be a martingale sequence such that |Xt − Xt−1 | ≤ c for every t, almost surely. Then for any
x > 0,
Xt − X0 tx2
P( ≥ x) ≤ 2e− 2c
t
Xt
As a result, t → 0 almost surely.

Backwards (reverse) martingales and decreasing information

An important class of martingales is known as backward martingales. A sequence of increasing σ-fields with a negative
time index,
· · · ⊂ F−n ⊂ F−n+1 ⊂ · · · ⊂ F−2 ⊂ F−1 ⊂ F0 ,
is called a reverse filtration. Note that here information is decreasing as n → −∞. Mn is called a backwards martingale
with respect to the reverse filtration if (i) E[|M−n |] < ∞, (ii) E[M−n+1 |F−n ] = M−n for all n ∈ Z.
Using similar arguments as those in the proof of the martingale convergence theorem studied earlier, we can arrive at the
following:

Theorem 4.1.16 Let (M−n , F−n ) be a backwards martingale sequence. Then, limn→−∞ Mn =: M−∞ = E[M0 | ∩∞
n=0
F−n ] almost surely and also in L1 .

4.2 Stability of Markov Chains: Foster-Lyapunov Techniques

Via martingale theory, a Markov chain’s stability can be characterized by drift conditions, as we discuss below in detail.
4.2 Stability of Markov Chains: Foster-Lyapunov Techniques 61

4.2.1 Criterion for invariance (existence of invariant probability measures) and positive Harris recurrence

Theorem 4.2.1 [Foster-Lyapunov for Positive Harris Recurrence] [232] Let S be a petite set, b ∈ R, ϵ > 0, and V : X →
R+ . If the following is satisfied for all x ∈ X:
Z
E[V (xt+1 )|xt = x] = P (x, dy)V (y) ≤ V (x) − ϵ + b1{x∈S} , (4.10)
X

then the chain is positive Harris recurrent (and thus a unique invariant probability measure π exists).

Proof. We will first assume that S is such that supx∈S V (x) < ∞. Define M̄0 := V (x0 ), and for t ≥ 1
t−1
X
M̄t := V (xt ) − (−ϵ + b1{xi ∈S} )
i=0

We have that
E[M̄(t+1) |xs , s ≤ t] ≤ M̄t , ∀t ≥ 0.
It follows from (4.10) that Ex [|M̄t |] ≤ ∞ for all t (by an application of the monotone convergence theorem applied
inductively: suppose that E[V (xt )] < ∞; then first show that E[E[min(N1 , V (xt+1 ))|xt ]] ≤ E[V (xt )] + b, then the take
the limit as N1 → ∞ to conclude that E[V (xt+1 )] < ∞ as well) and thus, {M̄t } is a supermartingale sequence with
respect to the natural filtration Ft = σ(x0 , · · · , xt ). Now, define a stopping time: τ N := min(τ, N ), where τ = min{i >
0 : xi ∈ S}. Note that the stopping time τ N is bounded. Hence, we have, by the martingale optional sampling theorem

E[M̄τ N |x0 ] ≤ M0 .

Thus, we obtain
N N
τX −1 τX −1
ϵEx0 [ 1] ≤ V (x0 ) + bEx0 [ 1{xi ∈S} ]
i=0 i=0

Thus,
ϵEx0 [τ N − 1 + 1] ≤ V (x0 ) + b,
and by the monotone convergence theorem (and that P (τ < ∞) = 1 by the uniform bound on the expectation below),

V (x0 ) + b
lim Ex0 [τ N ] = Ex0 [τ ] ≤ . (4.11)
N →∞ ϵ
Note that, the first equality above is a consequence of the drift criterion:

V (x0 ) + b
≥ lim Ex0 [τ N ] ≥ lim sup(N Px0 (τ ≥ N ) + Ex0 [τ 1{N >τ } ]) ≥ lim sup N Px0 (τ ≥ N ),
ϵ N →∞ N →∞ N →∞

implying that Px0 (τ ≥ N ) → 0 as N → ∞ and that Px0 (τS < ∞) = 1.) Now, if we had that

sup V (x) < ∞, (4.12)


x∈S

the proof would essentially be complete in view of Theorem 3.2.6. The fact that Ex [τ ] ≤ V (x)+b
ϵ < ∞ for any x ∈ X
leads to the Harris recurrence of the chain since this implies that Px (τ < ∞) = 1 for every x and petiteness implies that
the chain would be positive Harris recurrent [232, Proposition 9.1.7] (see also [89, Theorem 3.1]).
Typically, condition (4.12) is satisfied. However, the theorem statement does not impose this condition. Then, we proceed
with constructing another petite set on which (4.12) holds. Following [232, Chapter 11], define for some l ∈ Z+

VS (l) = {x ∈ S : V (x) ≤ l}
62 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

We will show that B := VS (l) is itself a petite set which is recurrent and satisfies the uniform finite-mean-return property. It
can be shown that, without any loss S is petite for some measure ν with ν(S) > 0, 1 and thus by a continuity of probability
argument, for sufficiently large l, we also have that ν(B) > 0. Again, since S is petite for measure ν, we have that

Ka (x, B) ≥ 1{x∈S} ν(B), x ∈ X,

a(i)P i (x, B), and hence


P
where Ka (x, B) = i∈N

1
1{x∈S} ≤ Ka (x, B)
ν(B)

Now, for x ∈ B,
B −1
τX B −1
τX
1
Ex [τB ] ≤ V (x) + bEx [ 1{xk ∈S} ] ≤ V (x) + bEx [ Ka (xk , B)] (4.13)
ν(B)
k=0 k=0
B −1
τX B −1 X
τX
1 1
= V (x) + b Ex [ Ka (xk , B)] = V (x) + b Ex [ a(i)P i (xk , B)] (4.14)
ν(B) ν(B) i
k=0 k=0
B τ −1
1 X X
= V (x) + b a(i)Ex [ 1{xk+i ∈B} ]
ν(B) i
k=0
1 X
≤ V (x) + b a(i)(1 + i), (4.15)
ν(B) i

P since at most once the process can hit B between 0 and τB − 1. Now, the petiteness measure can be
where (4.15) follows
adjusted such that i ai i < ∞ (by Theorem 3.2.3 or [232, Proposition 5.5.6]), leading to the result that

1 X
sup Ex [τB ] ≤ sup V (x) + b a(i)(1 + i) < ∞.
x∈B x∈B ν(B) i

Finally, since S is petite, so is B and it can be shown that Px (τB < ∞) = 1 for all x ∈ X. This concludes the proof. ⋄

Remark 4.2. Note that irreducibility of the Markov chain is not imposed a priori building on [232, Proposition 9.1.8]
or [89, Theorem 3.1], the drift criterion and the small/petite nature of the set leads to an irreducible Markov chain (possibly
defined on a proper subset of X).

Remark 4.3. Meyn and Tweedie [232, Theorem 13.0.1] show that under the hypotheses of Theorem 4.2.1, together with
aperiodicity, it also follows that for any initial state x ∈ X,

lim sup |P n (x, B) − π(B)| = 0,


n→∞ B∈B(X)

that is P n (x, · · · ) converges to π in total variation, for every x ∈ X. This follows from a coupling argument, to be discussed
further in the chapter.

Exercise 4.2.1 Consider a queuing system with

Qt+1 = max(Qt + At − N 1Qt ≥N , 0)

where At is an i.i.d. Poisson arrival process with rate λ so that


1
To see this, first take some set B with ν(B) > 0 and also with V (x) < M for all x ∈ B for some M large enough (such a set B
exists since the set {x : V (x) < ∞} is absorbing and then by a continuity of probability argument) and then consider the transitions
from B to S using the uniform probability of transitions by some time m large enough, by Markov’s inequality given (4.11), which then
provides a positive lower bound on transitions under a sampling distribution on Z+ from any x ∈ S to S via first visiting B and then
coming back to S.
4.2 Stability of Markov Chains: Foster-Lyapunov Techniques 63

λm
P (At = m) = e−λ , m ∈ Z+
m!
Suppose that N > λ. Show that Qt is positive Harris recurrent.

Remark 4.4. We note that if xt is aperiodic and irreducible and such that for some small set A we have supx∈A E[min(t >
0 : xt ∈ A)|x0 = x] < ∞, then the sampled chain {xkm } is such that supx∈A E[min(km > 0 : xkm ∈ A)|x0 = x] < ∞,
and the split chain discussion in Section 3.2.1 applies (See [232, Theorem 11.3.14]). The argument for this builds on the
fact that, with σC = min(k ≥ 0 : xk ∈ C), V (x) := 1+Ex [σC ], it follows that E[V (xt+1 )|xt = x] ≤ V (x)−1+b1{x∈C}
and iterating the expectation m times we obtain that
m−1
X
E[V (xt+m )|xt = x] ≤ V (x) − m + bEx [ 1{xk ∈C} ]. (4.16)
k=0

Pm−1
By [232], it follows that Ex [ k=0 1{xk ∈C} ] ≤ m1{x∈Cϵ } + mϵ for some petite set Cϵ and ϵ > 0 (this follows from the
observation that {x : P k (x, C) ≥ ϵ} will be included in the petite set for at least one k with 1 ≤ k ≤ m − 1 and the
complement of these sets {x : P k (x, C) < ϵ} will contribute to an upper bound of mϵ in (4.16)). This set is petite also
for the sampled chain (see Lemma 4.2.1). As a result, we have a drift condition for the m-skeleton, the return time for an
artificial atom constructed through the split chain is finite and hence an invariant probability measure for the m-skeleton,
and thus by (3.23), an invariant probability measure for the original chain exists. ⋄
In the following, we relax the existence of a petite set or irreducibility, but impose that the space is locally compact (and
not just Polish or standard Borel). This builds on [232, Theorem 12.3.4] or [167, Theorem 7.2.4].

Theorem 4.2.2 If the Markov chain is weak Feller, the space is locally compact, and S is compact; under (4.10), there
exists an invariant probability measure.

Proof. Iterating (4.10) we obtain that, with


n−1
1 X
P (n) (x, S) := Ex [ 1{xk ∈S} ],
n
k=0

we arrive at
ϵ
lim inf P (n) (x, S) ≥ .
n→∞ b
The result then follows from Theorem 3.3.2. ⋄
There are other versions of Foster-Lyapunov criteria, as we discuss in the following.

4.2.2 Criterion for finite expectations

Theorem 4.2.3 [Comparison Theorem] [232, Theorem 14.2.2] Let V : X → R+ , f, g : X → R+ . Let {xn } be a Markov
chain on X. If the following is satisfied:
Z
P (x, dy)V (y) ≤ V (x) − f (x) + g(x), x ∈ X,
X

then, for any stopping time τ with P (τ < ∞) = 1, it follows that


τ
X −1 τ
X −1
E[ f (xt )] ≤ V (x0 ) + E[ g(xt )]
t=0 t=0

Proof. As in Theorem 4.2.1, define M̄0 := V (x0 ), and for t ≥ 1


64 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics
t−1
X
M̄t := V (xt ) + (f (xi ) − g(xi )).
i=0

It follows that
E[M̄(t+1) |xs , s ≤ t] ≤ M̄t ,
∀t ≥ 0.
k−1
Now, define a stopping time: τ N = min(τ, min(k > 0 : k + V (xk ) + i=0 f (xk ) + g(xk ) ≥ N )). Note that the stopping
P
time τ N is bounded. It then follows that (through defining a supermartingale: Mt := M̄min(t,τ N ) ), and by the martingale
optional sampling theorem:
E[Mτ N |x0 ] ≤ M0 = V (x0 ).
Hence, we obtain
N
 τX −1 
E V (xτ N ) + (f (xi ) − g(xi )) x0 ≤ M̄0 = V (x0 ),
i=0

and thus by the fact that the terms inside the expectations are separately integrable, we have that
N N
τX −1 τX −1
E[ f (xi )|x0 ] ≤ M̄0 = V (x0 ) + E[ g(xi )|x0 ] − E[V (xτ N )|x0 ].
i=0 i=0

Now, since each of the terms in the expectations is positive, and that E[V (xτ N )|x0 ] ≥ 0, the monotone convergence
theorem implies the desired result. ⋄
Theorem 4.2.3 above also allows for the computation of useful bounds. For example if g(x) = b1{x∈A} , then one obtains
Pτ −1
R E[ t=0 f (xt )] ≤ V (x0 ) + b. In view of the invariant measure properties, if f (x) ≥ 1, this provides a bound on
that
π(dx)f (x), as we note next.

Theorem 4.2.4 [Criterion for finite expectations] [232] Let S be a petite set, b ∈ R+ and V : X → R+ , f : X → [ϵ, ∞)
for some ϵ > 0. Let {xn } be a Markov chain on X.

(i) If the following is satisfied:


Z
P (x, dy)V (y) ≤ V (x) − f (x) + b1{x∈S} , x ∈ X, (4.17)
X

then for every x0 = z ∈ X,


T −1 Z
1 X
lim f (xt ) = µ(dx)f (x) ≤ b, (4.18)
T →∞ T
t=0

almost surely, where µ is the invariant probability measure on X.


(ii) If {xt } is positive Harris recurrent, even if f : X → R+ (and not necessarily f : X → [ϵ, ∞) for some ϵ > 0) and
S = X itself (that is, with no indicator function), (4.17) implies (4.18).

That under Theorem 4.2.4, the process is a positive Harris recurrent Markov chain is a consequence of Theorem 4.2.1. The
proof of Theorem 4.2.4 will then build on the following result and the ergodicity of a positive Harris recurrent Markov
chain.

Theorem 4.2.5 Let (4.17) hold (but with not necessarily an irreducibility assumption), or the following more relaxed form
hold:
Z
P (x, dy)V (y) ≤ V (x) − f (x) + b, x∈X (4.19)
X
R
Under every invariant probability measure π, π(dx)f (x) ≤ b.
4.2 Stability of Markov Chains: Foster-Lyapunov Techniques 65

Proof. By Theorem 4.2.3, with taking T to be a deterministic stopping time, for any initial condition x0 = z
T −1  
1 X 1
lim sup Ez [ f (xk )] ≤ lim sup V (z) + bT = b. (4.20)
T →∞ T T →∞ T
k=0

Now, suppose that π is any invariant probability


R measure. Fix N < ∞, let fN = min(N, f ), and apply Fatou’s Lemma as
follows, where we use the notation π(f ) = π(dx)f (x),

 1 n−1
X 
π(fN ) = lim sup π P t fN
n→∞ n t=0
n−1
 1 X 
≤ π lim sup P t fN ≤ b.
n→∞ n t=0

Fatou’s Lemma is justified to obtain the first inequality, because fN is bounded. The monotone convergence theorem, with
taking N → ∞, then gives π(f ) ≤ b. ⋄

Remark 4.5. We note that where the system starts from, or what the initial distribution is on x0 is, affects the convergence
properties of
T −1
1 X
E[ f (xk )].
T
k=0

See Section 3.4.1 for a detailed [Link] particular, it is not necessarily the case that for every initial measure con-
vergence of expected normalized values to π(dx)f (x) holds. Furthermore, sample paths and expectations have slightly
different convergence characteristics for unbounded f .

4.2.3 Criterion for recurrence

Theorem 4.2.6 (Foster-Lyapunov for Recurrence) Let S be a compact set, b < ∞, and V be an inf-compact functional
from X → R+ such that for all α ∈ R+ {x : V (x) ≤ α} is compact (note: this implies that lim||x||→∞ V (x) = ∞ if
X = Rd for some d ∈ N). Let the following be satisfied for the Markov chain {xk }:
Z
P (x, dy)V (y) ≤ V (x) + b1{x∈S} , ∀x ∈ X , (4.21)
X

Furthermore, with τS = min(t > 0 : xt ∈ S), and τBN = min(t > 0 : xt ∈ BN ) where BN = {z : V (z) ≥ N }, if we
have that Px (min(τS , τBN ) = ∞) = 0 for every x ∈ X and N ∈ N, it must be that

Px (τS < ∞) = 1

for all x ∈ X

Proof. Define two stopping times: Let τS = min(t > 0 : xt ∈ S) and τBN = min(t > 0 : xt ∈ BN ) where BN =
{z : V (z) ≥ N } with N ≥ V (x) where x0 = x. Note that V (xt ) is bounded until τ N := min(τS , τBN ) and until this
time E[V (xt+1 )|Ft ] ≤ V (xt ). By assumption τ N = min(τS , τBN ) < ∞ with probability 1. Define Mt = V (xmin(t,τ N ) ),
which is a supermartingale sequence uniformly bounded (see Exercise 4.5.3). It follows then that a variation of the optional
sampling theorem (see Theorem 4.1.14) applies so that

Ex [Mτ N ] = Ex [V (xmin(τS ,τBN ) )] ≤ V (x)

Without any loss take x ∈


/ S (for otherwise, in the next time stage x1 we can make the argument replacing x0 with x1 ) and
there exists N large enough so that x ∈
/ (S ∪ BN ). Now, for x ∈ / (S ∪ BN ), since when exiting into BN the minimum
value of the Lyapunov function is N :
66 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

V (x) ≥ Ex [V (xmin(τS ,τBN ) )] ≥ Px (τBN < τS )N + Px (τBN ≥ τS )M,

for some finite non-negative M := inf x∈S V (x).


Hence,
V (x)
Px (τBN < τS ) ≤
.
N
We also have that P (min(τS , τBN ) = ∞) = 0. As a consequence, we have that

Px (τS = ∞) ≤ P (τBN < τS ) ≤ V (x)/N

and taking the limit as N → ∞, Px (τS = ∞) = 0. ⋄

Remark 4.6. If S is further petite, then once the petite set is visited, any other set with a positive measure (under an
irreducibility measure, since the petiteness measure can be used to construct an irreducibility measure) is visited with
probability 1 infinitely often and hence the chain is Harris recurrent. ⋄

Exercise 4.2.2 Show that the random walk on Z is Harris recurrent.

4.2.4 Criterion for transience

Criteria for transience is somewhat more difficult to establish. One convenient way is to construct a stopping time sequence
and show that the state does not come back to some set infinitely often. We state the following.

Theorem 4.2.7 (Criterion for Transience) [232], [153] Let V : X → R+ . If there exists a set A such that E[V (xt+1 )|xt =
x] ≤ V (x) for all x ∈
/ A and ∃x̄ ∈/ A such that V (x̄) < inf z∈A V (z), then {xt } is not recurrent, in the sense that
Px̄ (τA < ∞) < 1.

Proof. Let x = x̄. Proof follows from observing that


Z Z
V (x) ≥ V (y)P (x, dy) ≥ ( inf V (z))P (x, A) + V (y)P (x, dy) ≥ ( inf V (z))P (x, A)
y z∈A y ∈A
/ z∈A

It thus follows that


V (x)
P (τA < 2) = P (x, A) ≤
(inf z∈A V (z))
Likewise,
Z
V (x̄) ≥ V (y)P (x̄, dy)
X
Z Z
≥ ( inf V (z))P (x̄, A) + ( V (s)P (y, ds))P (x̄, dy)
z∈A y ∈A
/ X
Z  Z 
≥ ( inf V (z))P (x̄, A) + P (x̄, dy) ( inf V (s))P (y, A) + V (s)P (y, ds)
z∈A y ∈A
/ s∈A s∈A
/
Z
≥ ( inf V (z))P (x̄, A) + P (x̄, dy)(( inf V (s))P (y, A))
z∈A y ∈A
/ s∈A
 Z 
= ( inf V (z)) P (x̄, A) + P (x̄, dy)P (y, A) . (4.22)
z∈A y ∈A
/
R R
Thus, noting that P ({ω : τA (ω) < 3}) = A
P (x̄, dy) + y ∈A
/
P (x̄, dy)P (y, A), we observe:

V (x̄)
Px̄ (τA < 3) ≤ .
(inf z∈A V (z))
4.2 Stability of Markov Chains: Foster-Lyapunov Techniques 67
V (x̄)
Thus, this follows for any n: Px̄ (τA < n) ≤ (inf z∈A V (z)) < 1. Continuity of probability measures (by defining: Bn =
{ω : τA < n} and observing Bn ⊂ Bn+1 and that limn P (τA < n) = P (∪n Bn ) = P (τA < ∞) < 1) now leads to
Px̄ (τA < ∞) < 1. ⋄
Observe the striking difference with the inf-compactness condition leading to recurrence and the condition above, leading
to non-recurrence.

4.2.5 Criterion for almost sure convergence to equilibrium

The following build on stochastic stability theorems due to Khasminskii [192] and Kushner [203].

Theorem 4.2.8 (i) Let xn be Markov so that for some V : X → R+ and k : X → R+ , we have that

E[V (xn+1 )|xn = x] ≤ V (x) − k(x), x ∈ X.

Then, k(xn ) → 0 with probability 1.


(ii) Let Sλ := {x : V (x) ≤ λ}, and suppose that

E[V (xn+1 )|xn = x] ≤ V (x) − k(x), x ∈ Sλ .

If x0 ∈ Sλ , then,
 
Px 0 sup V (xn ) ≥ λ ≤ V (x0 )/λ. (4.23)
0≤n<∞

Hence, the paths remain in Qλ with probability at least 1− V (x 0)


λ . Furthermore, for paths that remain in Qλ , k(xn ) → 0
with probability 1.
(iii)Suppose that for each γ > 0, there exists δ > 0 so that k(x) ≥ δ for |x| ≥ γ and that k(0) = 0. Then, the origin is
globally asymptotically stable with probability 1, that is, limn→∞ xn = 0 almost surely.
(iv) Suppose that for some increasing function c : X → R+ with c(0) = 0 and c(x) > 0 for all x ̸= 0, we have that
c(|x|) ≤ V (x) and for some α > 0

E[V (xn+1 )|xn = x] ≤ V (x) − αV (x), x ∈ X.

Then, the system is exponentially asymptotically stable in the sense that:

V (x0 )(1 − α)N


Px0 { sup V (xn ) ≥ λ} ≤ .
N ≤n<∞ λ

Proof.
Pn−1 P∞
(i) By argumentsP presented earlier, it follows that 0 ≤ Ex0 [V (xn )] ≤ V (x0 )−Ex0 [ m=0 k(xm )]. Thus, Ex0 [ m=0 k(xm )] <

∞. But then, m=0 k(xm ) < ∞ almost surely and thus k(xm ) → 0 almost surely.
Pn−1
(ii) Define M0 = V (x0 ) and for n > 0 : Mn = V (xn ) − m=0 k(xm ), which is a supermartingale sequence with
respect to the natural filtration. Let us stop the process xn on first leaving Sλ where Sλc = X \ Sλ . Then, the stopped
process Mmin(t,τSc ) is also a supermartingale process, where the drift equation holds with k(x) = 0 for x ∈ / Sλ and if
λ
τSλc = ∞, we have that k(xm ) → 0.
On the other hand, the bound in (4.23) builds essentially on the proof of Doob’s maximal inequality Theorem 4.1.7,
which notes that for a non-negative supermartingale Rn , for all λ > 0,
68 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

R0
P ( max Rn ≥ λ) ≤
0≤n<∞ λ
With Mn the super-martingale sequence defined as before, we have that
τS c −1
 λ 
X M0
Px0 (V (xτSc ) ≥ λ) ≤ Px0 V (xτSc ) − k(xm ) ≥ λ ≤
λ λ
m=0
λ

(iii)By (i), we have that k(xn ) → 0; the hypothesis then implies that xn → 0.
V (xn )
(iv)Observe first that Mn := (1−α)n is also a supermartingale. Apply Doob’s maximal inequality (Theorem 4.1.7) as
follows:
V (xn )(1 − α)n V (xn ) λ E[MN ](1 − α)N M0 (1 − α)N
P ( sup ≥ λ) ≤ P ( sup ≥ ) ≤ ≤ .
N ≤n<∞ (1 − α)n N ≤n<∞ (1 − α)
n (1 − α)N λ λ

4.2.6 State dependent drift criteria: Deterministic and random-time

In many applications, a drift term (e.g. by a controller) can be applied on a system only intermittently.

Theorem 4.2.9 [350] Suppose that {xt } is a φ-irreducible and aperiodic Markov chain. Suppose moreover that there are
functions V : X → (0, ∞), δ : X → [1, ∞), f : X → [1, ∞), a small set C on which V is bounded, and a constant b ∈ R,
such that
E[V (xτi+1 ) | Fτi ] ≤ V (xτi ) − δ(xτi ) + b1C (xτi )
X−1
 τi+1 
(4.24)
E f (xk ) Fτi ≤ δ(xτi ) , i ≥ 0.
k=τi

Then the following hold:

(i) {xt } is positive Harris recurrent, with unique invariant distribution π


R
(ii) π(f ) := f (x) π(dx) < ∞.
(iii) For any function g that is bounded by f , in the sense that supx |g(x)|/f (x) < ∞, we have convergence of
moments in the mean, and the strong law of large numbers holds:

lim Ex [g(xt )] = π(g)


t→∞
N −1
1 X
lim g(xt ) = π(g) a.s. , x ∈ X
N →∞ N
t=0

By taking f (x) = 1 for all x ∈ X, we obtain the following corollary to Theorem 4.2.9.

Corollary 4.2.1 [350] Suppose that X is a φ-irreducible Markov chain. Suppose moreover that there is a function V :
X → (0, ∞), a petite set C on which V is bounded,, and a constant b ∈ R, such that the following hold:

E[V (xτz+1 ) | Fτz ] ≤ V (xτz ) − 1 + b1{xτz ∈C}


sup E[τz+1 − τz | Fτz ] < ∞. (4.25)
z≥0

Then X is positive Harris recurrent. ⋄


4.2 Stability of Markov Chains: Foster-Lyapunov Techniques 69

The above extend the deterministic state-dependent results presented in [232], [233]: Let τz , z ≥ 0 be a sequence of
stopping times, measurable on a filtration, possible generated by the state process.
Without the irreducibility condition, if the chain is weak Feller, if (4.10) holds with S compact, then there exists at least
one invariant probability measure as discussed in Section 3.3.1.

Theorem 4.2.10 [350] Suppose that X is a Feller Markov chain, not necessarily φ-irreducible. If (4.24) holds with C
compact then there exists at least one invariant probability measure. Moreover, there exists c < ∞ such that, under any
invariant probability measure π, Z
Eπ [f (x)] = π(dx)f (x) ≤ c. (4.26)
X

Petite sets and sampling

Unfortunately the techniques we reviewed earlier that rely on petite sets become unavailable in the random time drift setting
considered in Section 4.2.6 as a petite set C for {xn } is not necessarily petite for {xτn }. Some of the discussion in this
section is due to Zurkowski et. al. [354]

Lemma 4.2.1 [354] Suppose {xt } is an aperiodic and irreducible Markov chain. If there exists sequence of stopping times
{τn } independent of {xt }, then any C that is small for {xt } is petite for {xτn }.

Proof. Since C is petite, it is small by Theorem 3.2.4 for some m. Let C be (m, δ, ν)-small for {xt }.

X ∞
X Z
P τ1 (x, · ) = P (τ1 = k)P k (x, · ) ≥ P (τ1 = k) P m (x, dy)P k−m (y, · )
k=1 k=m

X Z
≥ P (τ1 = k) 1C (x)δν(dy)P k−m (y, · ) (4.27)
k=m


P (τ1 = k)P k−m (y, · ), we have that C is
R P
which is a well defined measure. Therefore defining κ( · ) = ν(dy)
k=m
(1, δ, κ)-small for {xτn }. ⋄
Thus, one can relax the condition that V is bounded on C in Theorem 4.2.9, if the sampling times are deterministic. Another
condition is when the sampling instances are hitting times to a set which contains C [354].

4.2.7 Convergence Rates to Equilibrium

In addition to obtaining bounds on the rate of convergence through Dobrushin’s coefficient studied earlier, a more relaxed
and often more general approach is via Foster-Lyapunov drift conditions and an associated coupling analysis.
Regularity and ergodicity are concepts closely related through the work of Meyn and Tweedie [232], [236] and Tuominen
and Tweedie [308].

Definition 4.2.1 A set A ∈ B(X ) is called (f, r)-regular if

B −1
τX
sup Ex [ r(k)f (xk )] < ∞
x∈A
k=0

for all B ∈ B + (X ). A finite measure ν on B(X ) is called (f, r)-regular if


70 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

B −1
τX
Eν [ r(k)f (xk )] < ∞
k=0

for all B ∈ B + (X ), and a point x is called (f, r)-regular if the measure δx is (f, r)-regular.

This leads to a lemma relating regular distributions to regular atoms.

Lemma 4.2.2 If a Markov chain {xt } has an atom α ∈ B + (X ) and an (f, r)-regular distribution λ, then α is an (f, r)-
regular set.

Definition 4.2.2 (f -norm) For a function f : X → [1, ∞) the f -norm of a measure µ defined on (X, B(X )) is given by
Z
∥µ∥f = sup | µ(dx)g(x)|.
g≤f

The total variation norm is the f -norm when f = 1, denoted by ∥.∥T V .

Definition 4.2.3 A Markov chain {xt } with invariant distribution π is (f, r)-ergodic if

r(n)∥P n (x, · ) − π( · )∥f → 0 as n → ∞ for all x ∈ X. (4.28)

If (4.28) is satisfied for a geometric r (so that r(n) = M ζ n for some ζ > 1, M < ∞) and f = 1 then the Markov chain
{xt } is called geometrically ergodic.

Coupling inequality and moments of return times to a small set The main idea behind the coupling inequality is to
bound the total variation distance between the distributions of two random variables by the probability they are different
Let X and Y be two jointly distributed random variables on a space X with distributions µx , µy respectively. Then we can
bound the total variation between the distributions by the probability the two variables are not equal.

∥µx − µy ∥T V = sup|µx (A) − µy (A)|


A
= sup|P (X ∈ A, X = Y ) + P (X ∈ A, X ̸= Y )
A
− P (Y ∈ A, X = Y ) − P (Y ∈ A, X ̸= Y )|
≤ sup|P (X ∈ A, X ̸= Y ) − P (Y ∈ A, X ̸= Y )|
A
≤P (X ̸= Y )

The coupling inequality is useful in discussions of ergodicity when used in conjunction with parallel Markov chains.
Later, we will see that the coupling inequality is also useful to establish the existence of optimal solutions to average cost
optimization problems.
One creates two Markov chains having the same one-step transition probabilities. Let {xn } and {x′n } be two Markov
chains that have probability transition kernel P (x, ·), and let C be an (m, δ, ν)-small set. We use the coupling construction
provided by Roberts and Rosenthal [264], building on the splitting technique presented in Section 3.2.1.

Let x0 = x and x′0 ∼ π where π is the invariant probability measure for both Markov chains.

(1) If xn = x′n then xn+1 = x′n+1 ∼ P (xn , ·)

(2) Else, if (xn , x′n ) ∈ C × C then with probability δ, xn+m = x′n+m ∼ ν(·) with probability 1 − δ then
4.2 Stability of Markov Chains: Foster-Lyapunov Techniques 71

independently
1 m
xn+m ∼ 1−δ (P (xn , ·) − δν(·))

1
x′n+m ∼ m ′
1−δ (P (xn , ·) − δν(·))

(3) Else, independently xn+m ∼ P m (xn , ·) and x′n+m ∼ P m (x′n , ·).

The in-between states xn+1 , ...xn+m−1 , x′n+1 , ...x′n+m−1 are distributed conditionally given xn , xn+m ,x′n , x′n+m .
By the Coupling Inequality and the previous discussion with Nummelin’s Splitting technique in Section 3.2.1 we have
∥P n (x, ·) − π(·)∥T V ≤ P (xn ̸= x′n ).

Remark 4.7. Through the coupling inequality one can show that π0 P n → π in total variation. Furthermore, if The-
orem 4.2.4 holds, one can also show that with some further analysis if the initial condition is a fixed determinis-
P n (x, dz) − π(dz) f (z) → 0, where f is not necessarily bounded. This does not imply, however,
R
tic state
 
n
R
(π0 P )(dz) − π(dz) f (z) → 0 for a random initial condition. A sufficient condition for the latter to occur is that
R
π0 (dz)V (z) < ∞ provided that Theorem 4.2.4 holds (see Theorem 14.3.5 in [232]).

Rates of convergence: Geometric ergodicity

In this section, following [232] and [264], we review results stating that a strong type of ergodicity, geometric ergodicity,
follows from a simple drift condition. An irreducible Markov chain is said to satisfy the univariate drift condition if there
are constants λ ∈ (0, 1) and b < ∞, along with a function V : X → [1, ∞), and a small set C such that

P V ≤ λV + b1C . (4.29)

Theorem 4.2.11 [264, Theorem 9] Suppose {xt } is an aperiodic, irreducible Markov chain with invariant distribution π.
Suppose C is a (1, ϵ, ν)-small set and V : X → [1, ∞) satisfies the univariate drift condition with constants λ ∈ (0, 1) and
b < ∞. Then {xt } is geometrically ergodic.

That geometric ergodicity follows from the univariate drift condition with a small set C is proven by Roberts and Rosenthal
by using the coupling inequality to bound the T V -norm, but an alternate proof is given by Meyn and Tweedie [232]
resulting in the following theorem.

Theorem 4.2.12 [232, Theorem 15.0.1] Suppose {xt } is an aperiodic and irreducible Markov chain. Then the following
are equivalent:
(i) Ex [τB ] < ∞ for all x ∈ X, B ∈ B + (X), the invariant distribution π of {xt } exists and there exists a petite set C,
constants γ < 1, M > 0 such that for all x ∈ C

|P n (x, C) − π(C)| < M γ n .

(ii) For a petite set C and for some κ > 1


sup Ex [κτC ] < ∞.
x∈C

(iii) For a petite set C, constants b > 0 λ ∈ (0, 1), and a function V : X → [1, ∞] (finite for some x) such that

P V ≤ λV + b1C .

Any of the conditions imply that there exists r > 1, R < ∞ such that for any x
72 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

X
rn ∥P n (x, · ) − π( · )∥V ≤ RV (x).
n=0

We note that if (iii) above holds, (ii) holds for for all κ ∈ (1, λ−1 ).
We now show that under (4.29), Theorem 4.2.12 (ii) holds. If (4.29) holds, the sequence {Mn } is supermartingale (with
respect to the natural filtration), where
n−1
X
Mn = λ−n V (xn ) − b1C (xk )λ−(k+1) ,
k=0

with M0 = V (x0 ). Then, with (4.29), defining τBN = min{N, τB } for B ∈ B + (X) gives, by Doob’s optional sampling
theorem,
N
  B −1
 τX 
N
−τB
Ex λ V (xτBN ) ≤ V (x) + Ex b1C (xn )λ−(n+1) (4.30)
n=0

for any B ∈ B + (X), and N ∈ Z+ .


Since V is bounded above on C, we have that C ⊂ {V ≤ L1 } for some L1 and thus,
 
N
−τC
sup Ex λ V (xτCN ) ≤ L1 + λ−1 b.
x∈C

and by the monotone convergence theorem, and the fact that V is bounded from below by 1 everywhere and bounded from
above on C,  
sup Ex λ−τC ≤ L1 (L1 + λ−1 b).
x∈C

Using the coupling inequality, Roberts and Rosenthal [264] prove that geometric ergodicity follows from the univariate drift
condition. They show that under mild conditions [264, Prop. 11], the univariate drift condition implies a drift condition for
the pair of Markov chains who will be coupled in the small set C × C:

Proposition 4.2.1 [264, Proposition 11] Suppose the univariate drift condition (4.29) is satisfied for V : X → [1, ∞) and
b
constants λ ∈ (0, 1) b < ∞ and small set C. Letting d = inf x∈C C V (x), if d > 1−λ − 1, then the bivariate drift condition
1 −1
is satisfied for h(x, y) = 2 (V (x) + V (y)) and α = λ + b/(d + 1) < 1; that is, we have the following condition.

h(x, y)
P̄ h(x, y) ≤ (x, y) ∈
/ C ×C
α
P̄ h(x, y) <∞ (x, y) ∈ C × C

where
Z Z
P̄ h(x, y) = h(z, w)P (x, dz)P (y, dw)
X X

But now if one applies Theorem 4.2.12 (ii), the desired coupling condition and hence the convergence rate result will
follow.
We also note that the univariate drift condition allows us to assume that V is bounded on C without any loss (see Lemma
14 of [264]).
4.2 Stability of Markov Chains: Foster-Lyapunov Techniques 73

Subgeometric ergodicity

Here, we review the class of subgeometric rate functions (see [153, Sec. 4], [88, Sec. 5], [232], [107], [308]).
Let Λ0 be the family of functions r : N → R>0 such that

r is non-decreasing, r(1) ≥ 2

and
log r(n)
↓ 0 as n → ∞
n
The second condition implies that for all r ∈ Λ0 if n > m > 0 then

n log r(n + m) ≤ n log r(n) + m log r(n) ≤ n log r(n) + n log r(m)

so that
r(m + n) ≤ r(m)r(n) for all m, n ∈ N. (4.31)

The class of subgeometric rate functions Λ defined in [308] is the class of sequences r for which there exists a sequence
r0 ∈ Λ0 such that
r(n) r(n)
0 < lim inf ≤ lim sup < ∞.
n→∞ r0 (n) n→∞ r0 (n)

The main theorem we cite on subgeometric rates of convergence is due to Tuominen and Tweedie [308].

Theorem 4.2.13 [308, Theorem 2.1] Suppose that {xt }t∈N is an irreducible and aperiodic Markov chain on state space X
with stationary transition probabilities given by P . Let f : X → [1, ∞) and r ∈ Λ be given. The following are equivalent:
(i) there exists a petite set C ∈ B(X) such that

C −1
τX
sup Ex [ r(k)f (xk )] < ∞
x∈C
k=0

(ii) there exists a sequence (Vn ) of functions Vn : X → [0, ∞], a petite set C ∈ B(X) and b ∈ R+ such that V0 is bounded
on C,
V0 (x) = ∞ ⇒ V1 (x) = ∞,
and
P Vn+1 ≤ Vn − r(n)f + br(n)1C , n∈N

(iii) there exists an (f, r)-regular set A ∈ B + (X).


(iv) there exists a full absorbing set S which can be covered by a countable number of (f, r)-regular sets.

Theorem 4.2.14 [308] If a Markov chain {xt } satisfies Theorem 4.2.13 for (f, r) then r(n)∥P n (x0 , ·) − π(·)∥f → 0 as
n increases

The conditions of Theorem 4.2.13 may be hard to check, especially (ii), comparing a sequence of Lyapunov functions {Vk }
at each time step. We briefly discuss the methods of Douc et al. [107] (see also Hairer [153]) that extend the subgeometric
ergodicity results and show how to construct subgeometric rates of ergodicity from a simpler drift condition. [107] assumes
that there exists a function V : X → [1, ∞], a concave monotone nondecreasing differentiable function ϕ : [1, ∞] →
(0, ∞], a set C ∈ B(X) and a constant b ∈ R such that

P V + ϕ◦V ≤ V + b1C . (4.32)


74 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

If an aperiodic and irreducible Markov chain {xt } satisfies the above with a petite set C, and if V (x0 ) < ∞, then it can
be shown that {xt } satisfies Theorem 4.2.13(ii). Therefore {xt } has invariant distribution π and is (ϕ◦V, 1)-ergodic so that
lim ∥P n (x, · ) − π( · )∥ϕ◦V = 0 for all x in the set {x : V (x) < ∞} of π-measure 1. The results by Douc et al. build then
n→∞
on trading off (ϕ ◦ V, 1) ergodicity for (1, rϕ )-ergodicity for some rate function rϕ , by carefully constructing the function
utilizing concavity; see Propositions 2.1 and 2.5 of [107] and Theorem 4.1(3) of [153].
To achieve ergodicity with a nontrivial rate and norm one can invoke a result involving the class of pairs of ultimately
non decreasing functions, defined in [107]. The class Y of pairs of ultimately non decreasing functions consists of pairs
Ψ1 , Ψ2 : X → [1, ∞) such that Ψ1 (x)Ψ2 (y) ≤ x + y and Ψi (x) → ∞ for one of i = 1, 2.

Proposition 4.2.2 Suppose {xt } is an aperiodic and irreducible Markov chain that is both (1, r)-ergodic and (f, 1)-ergodic
for some r ∈ Λ and f : X → [1, ∞). Suppose Ψ1 , Ψ2 : X → [1, ∞) are a pair of ultimately non decreasing functions. Then
{xt } is (Ψ1 ◦ f, Ψ2 ◦ r)-ergodic.

Therefore we can show that if (Ψ1 , Ψ2 ) ∈ Y and a Markov chain satisfies the condition (4.32), then it is (Ψ1 ◦ϕ◦V, Ψ2 ◦rϕ )-
ergodic.
Thus, we observe that the hitting times to a small set is an important random variable in characterizing not only the existence
of an invariant probability measure, but also how fast a Markov chain converges to equilibrium. Further results exist in the
literature to obtain more computable criteria for subgeometric rates of convergence, see e.g. [107].

Rates of convergence under random-time state-dependent drift criteria

The following result builds on and generalizes Theorem 2.1 in [350].

Theorem 4.2.15 [354] Let {xt } be an aperiodic and irreducible Markov chain with a small set C. Suppose there are
functions V : X → (0, ∞) with V bounded on C, f : X → [1, ∞), δ : X → [1, ∞), a constant b ∈ R, and r ∈ Λ such that
for a sequence of stopping times {τn }

E[V (xτn+1 ) | xτn ] ≤ V (xτn ) − δ(xτn ) + b1C (xτn )


X−1
 τn+1 
E f (xk )r(k) Fτn ≤ δ(xτn ). (4.33)
k=τn

Then {xt } satisfies Theorem 4.2.13 and is (f, r)-ergodic.

Further conditions and examples are available in [354].

4.3 Applications to Stochastic Learning Algorithms and Iterative Dynamics

In this section, we present a number of applications of martingale theory to the analysis of stochastic dynamics, which will
have applications to stochastic learning and reinforcement learning results to be studied later in the book.
We start with a convergence theorem useful in stochastic approximation]

Theorem 4.3.1 [243, p. 33, Exercise II-4] Let Xk , βk , Yk be three sequences of non-negative random variables defined on
a common probability space and Fk be a filtration so that all three random sequences are adapted to it. Suppose that

E[Xk+1 |Fk ] ≤ (1 + βk )Xk + Yk , k ∈ N.


P
Then,
P limit limn→∞ Xn exists and is finite with probability one conditioned on the event that k∈N βn < ∞ and
k∈N Yn < ∞.
4.3 Applications to Stochastic Learning Algorithms and Iterative Dynamics 75

Proof sketch. (i) Define


n−1
X
Mn = Xn′ − Ym′ ,
m=1

where Xn′ = Qn−1Xn and Yn′ = Qn Yn


. Define the stopping time:
(1+βs ) (1+βs )
s=1 s=1

n−1
X
τa = min(n : Ym′ > a).
m=1

(ii) Show first that a + Mmin(τa ,n) , n ∈ N is a positive supermartingale (iii) Thus, for any fixed a, a + Mmin(τa ,n)
P, n ∈ N
converges to a limit. For a given sample path (almost surely), taking a sufficiently large a, by the boundedness of k∈N Yn
show that τa = ∞ for sufficiently large P a for this given sample path. (iv) Then invoke
P the supermartingaleQ convergence
theorem 4.1.9. Finally, using the fact that k∈N Yk < ∞ and the implication (from k∈N βn < ∞) that n (1+βn ) < ∞
under the stated conditions, complete the proof. ⋄
This theorem is important for a large class of optimization problems (such as the convergence of stochastic gradient descent
algorithms) as well as stochastic approximation algorithms. For further reading on stochastic approximation methods,
see [208] and [34] and for a recent review [319]. We will use this result to establish the convergence of the celebrated
Q-learning algorithm in Theorem 9.1.1. A slight generalization of this result appears in [263].

Theorem 4.3.2 [Another convergence theorem useful in stochastic approximation and Q-Learning convergence analysis]
Let Xk , Yk , Zk be three sequences of non-negative random variables defined on a common probability space and Fk be a
filtration so that all three random sequences are adapted to it. Suppose that

E[Yk+1 |Fk ] ≤ Yk − Xk + Zk (4.34)


P P
and k Zk < ∞. Then, k Xk < ∞ and Yk converges to some random variable Y almost surely.

Proof sketch [341].: Apply Theorem 4.3.1 by noting first that E[Yk+1 |Fk ] ≤ Yk + Zk with βk = 0. This implies that
Pt−1
Yk converges. Now write Mt = Yt + m=1 Xm leading to PE[Mt+1 |Ft ] ≤ Mt + Zt . Applying Theorem 4.3.1 again, it
follows that Mt converges and since Yt converges, so does t Xt . ⋄
We now apply the above to an explicit iterative stochastic dynamics:

Theorem 4.3.3 [38, Corollary 4.1] Consider the following: Let rt be a scalar and

rt+1 = (1 − αt )rt + αt wt ,

αt2 < ∞, and the noise wt is so that E[wt |Ft−1 ] = 0 with


P P
where t αt = ∞, t

E[wt2 |Ft ] ≤ At ,

where At is possibly a random variable (thus, sample path dependent). If At is bounded with probability 1 (that is,
supt∈Z+ |At (ω)| < ∞ almost surely), then rt → 0 almost surely.

We remark that if the bounded random variable sequence At above was instead a fixed number, the proof would have been
slightly more direct.
Proof. Take Yt = rt2 and apply Theorem 4.3.2. ⋄
We end the section, with a final application2 :

2
Thanks to Prof. Jerome Le Ny (of Polytechnique Montreal).
76 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

Theorem 4.3.4 [Application in stochastic optimization: Stochastic gradient descent] Consider a convex function f : Rn →
R and denote the set of minima of f by X ∗ . We know from convex analysis that X ∗ contains, if non-empty, either a single
point, or is a convex set. Denote the subdifferential [69] of f at x, that is the set of subgradients of f at x, by ∂f (x) and let
dt be a random variable which is a noisy version of a sub gradient of f at xt at time t. A stochastic subgradient algorithm
is one with the form:

xk+1 = xk − γk dk+1 , x0 ∈ Rn , (4.35)

where γk is a sequence of non-negative step sizes. We have the following theorem:


Suppose that the set of minima X ∗ is non-empty and that the stochastic subgradients satisfy that

sup E[||dk+1 ||2 |Fk ] < K < ∞.


k

where Fk = σ(x0 , ds , s ≤ k) with the condition that

gk+1 = E[dk+1 |Fk ] ∈ ∂f (x).

γk2 < ∞. Then the sequence of iterates (4.35) converges almost surely to some element
P P
Moreover, k γk = ∞ and k
x∗ ∈ X ∗ .

Proof. y ∈ Rn , we have that due to the definition of a subgradient,


T
f (y) ≥ f (x) + gk+1 (y − xk ).(4.36)

Thus,

E[||xk+1 − y||2 |Fk ] = E[||xk − γk dk+1 − y||2 |Fk ]


= E[||xk − y||2 − 2γk (xk − y)T dk+1 + γk2 ||dk+1 ||2 |Fk ]
= E[||xk − y||2 |Fk ] − 2γk (xk − y)T E[dk+1 |Fk ] + Eγk2 ||dk+1 ||2 |Fk ]
≤ E[||xk − y||2 |Fk ] − 2γk (f (xk ) − f (y)) + γk2 K

where in the inequality we use (4.35). Now, let in the above y = x̄∗ ∈ X ∗ for some element in X ∗ . Then one obtains
through the comparison theorem (Theorem 4.2.3) that
X X
E[ γk (f (xk ) − f (y))] ≤ ||x0 − y||2 + γk2 K.
k k

In particular, since f (xk ) − f (y) ≥ 0, through the convergence theorem from the preceding exercises we have that almost
surely X
γk (f (xk ) − f (y)) < ∞.
k

Thus, f (xk ) → f (y). We now show that indeed xk → some particular element in X ∗ (and does not wander in the set). By
the convergence result in (4.34) we know that for any x∗ ∈ X ∗ , ||xk − x∗ || converges almost surely. This implies that xk
is bounded almost surely. Now, consider a countable dense subset {x1,∗ , · · · , xn,∗ , · · · } of X ∗ . It must be that ||xk − xi,∗ ||
converges for all i through the convergence theorem. On the other hand, since ||xk || is bounded, there exists a converging
subsequence for xkn . But the limit of each such subsequence must be identical for otherwise ||xkn − xi,∗ || would have
different limits. Thus, xkn must converge to one element in X ∗ . ⋄

4.4 Conclusion

This concludes our discussion on martingales, their applications to controlled Markov chains, ergodic theorems, as well
as stochastic iterative dynamics. We will revisit one more application of martingales while discussing the convex analytic
4.5 Exercises 77

approach to controlled Markov problems. Notably, we have observed that drift criteria are very powerful tools to establish
various forms of stochastic stability and instability.

4.5 Exercises

Exercise 4.5.1 Let X be an integrable random variable defined on (Ω, F, P ). Let G = {Ω, ∅}. Show that E[X|G] = E[X],
and if G = σ(X) then E[X|G] = X.

Exercise 4.5.2 Consider (4.5) and through this relation establish a sufficient condition on the martingale sequence Xn so
that the optimal sampling theorem would be applicable even if the stopping times in Theorem 4.1.5 would not necessarily
be bounded from above by a deterministic constant.

Exercise 4.5.3 A useful property of martingales is that a stopped martingale is a martingale. This is very useful for proving
stability results when one lives in a bounded set since the stopped martingale sequence will typically be uniformly bounded
(and hence the optional sampling theorem will be applicable without requiring a stopping time to be uniformly bounded).
Let τ be a stopping time that is finite almost surely. Let Xt , Ft be a martingale sequence. Define Mt = Xmin(t,τ ) . Show
that (Mt , Ft ) is a martingale sequence:
E[Mn+1 |Fn ] = Mn
Hint: Write Mn+1 = Mn + 1{τ >n} (Mn+1 − Mn ). Then, show that E[1{τ >n} (Mn+1 − Mn )|Fn ] = 1{τ >n} E[Mn+1 −
Mn |Fn ] = 0.

Exercise 4.5.4 a) Consider a Controlled Markov Chain with the following dynamics:

xt+1 = axt + but + wt ,

where wt is a zero-mean Gaussian noise with a finite variance, a, b ∈ R, b ̸= 0, are the system dynamics coefficients. One
controller policy which is admissible (that is, the policy at time t is measurable with respect to σ(x0 , x1 , . . . , xt ) and is a
mapping to R) is the following:
a + 0.5
ut = − xt .
b
Show that {xt }, under this policy, has a unique invariant probability measure.
b) Consider a similar setup to the one earlier, with b = 1:

xt+1 = axt + ut + wt ,

where wt is a zero-mean Gaussian noise with a finite variance, and a ∈ R is a known number.
This time, suppose, we would like to find a control policy such that there exists an invariant probability measure π for {xt }
and under this invariant probability measure
Eπ [x2 ] < ∞
Further, suppose we restrict the set of control policies to be linear, time-invariant; that is of the form u(xt ) = kxt for some
k ∈ R.
Find the set of all k values for which there exists an invariant probability measure that has a finite second moment.
Hint: Use Foster-Lyapunov criteria.

Exercise 4.5.5 Suppose that some price process {xt , t ∈ Z+ } is given by the following dynamics:

xt+1 = max(xt + wt , 0), t ∈ Z+ ,


78 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics

where {wt } is a sequence of independent and identically distributed {−1, 1}-valued random variables with mean w̄ > 0.
Furthermore, x0 ∈ Z+ , x0 > 0 is a given initial condition for the process.
Is the price process recurrent in the sense that, Px0 (τ0 < ∞) = 1, where τ0 = min(l > 0 : xl = 0)?

Exercise 4.5.6 Consider a queuing process, with i.i.d. Poisson arrivals and departures, with arrival mean µ and service
mean λ and suppose the process is such that when a customer leaves the queue, with probability p (independent of time) it
comes back to the queue. That is, the dynamics of the system satisfies:

Lt+1 = max(Lt + At − Nt + pt Nt , 0), t ∈ N.

where E[At ] = λ, E[Nt ] = µ and E[pt ] = p.


For what values of µ, λ is such a system stochastically stable? Prove your statement.

Exercise 4.5.7 Consider a two server-station network; where a router routes the incoming traffic, as is depicted in Figure
5.1.

Station 1

Station 2

Fig. 4.1

Let L1t , L2t denote the number of customers in stations 1 and 2 at time t. Let the dynamics be given by the following:

L1t+1 = max(L1t + ut At − Nt1 , 0), t ∈ N.

L2t+1 = max(L2t + (1 − ut )At − Nt2 , 0), t ∈ N.

Customers arrive according to an independent Bernoulli process, At , with mean λ. That is, P (At = 1) = λ and P (At =
0) = 1 − λ. Here ut ∈ [0, 1] is the router action.
Station 1 has a Bernoulli service process Nt1 with mean n1 , and Station 2 with n2 .
Suppose that a router decides to follow the following algorithm to decide on ut : If a customer arrives, the router simply
sends the incoming customer to the shortest queue.
Find sufficient conditions (on λ, n1 , n2 ) for this algorithm to lead to a stochastically stable system with invariant measure
π which satisfies Eπ [L1 + L2 ] < ∞.
Note: For this problem, we acknowledge the lecture notes of Prof. Bruce Hajek: ECE567 Communication Network Analysis,
University of Illinois at Urbana-Champaign [156].

Exercise 4.5.8 Consider the following two-server system:

x1t+1 = max(x1t + A1t − u1t , 0)


x2t+1 = max(x2t + A2t + u1t 1(u1t ≤x1t +A1t ) − u2t , 0), (4.37)

where 1(.) denotes the indicator function and A1t , A2t are independent and identically distributed (i.i.d.) random variables
with geometric distributions, that is, for i = 1, 2,
4.5 Exercises 79

P (Ait = k) = pi (1 − pi )k k ∈ {0, 1, 2, . . . , },

for some scalars p1 , p2 such that E[A1t ] = 1.5 and E[A2t ] = 1.


Suppose the control actions u1t , u2t are such that u1t + u2t ≤ 5 for all t ∈ Z+ and u1t , u2t ∈ Z+ . At any given time t, the
controller has to decide on u1t and u2t with knowing {x1s , x2s , s ≤ t} but not knowing A1t , A2t .
Is this server system stochastically stabilizable by some policy, that is, does there exist an invariant probability measure
under some control policy?
If your answer is positive, provide a control policy and show that there exists a unique invariant distribution.

Exercise 4.5.9 Let there be a single server, serving two queues; where the server serves the two queues adaptively in the
following sense. The dynamics of the two queues is expressed as follows:

Lit+1 = max(Lit + Ait − Nti , 0), i = 1, 2; t ∈ Z+

where Lit is the total number of arrivals which are still in the queue at time t and Ait is the number of customers that have
just arrived at time t.
We assume, for i = 1, 2, {Ait } has an independent and identical distribution (i.i.d.) which is Bernoulli so that P (Ait =
1) = λi = 1 − P (Ait = 0).
Suppose that the service process is given by:

Nt1 = 1{L1t ≥L2t } Nt2 = 1{L2t >L1t }

For what values of λ1 , λ2 is the system stochastically stable, in the sense of the existence of an invariant probability
measure.

Exercise 4.5.10 Let X be a real random variable with E[|X|] < ∞. Let Y0 , Y1 , Y2 , · · · be a sequence of random variables.
Let Fn be the σ-field generated by Y0 , Y1 , . . . , Yn . a) Is it the case that

lim E[X|Fn ]
n→∞

exists? b) Is it the case that


lim E[X|Fn ] = E[X|F∞ ],
n→∞

where F∞ := σ(Y1 , Y2 , · · · )

Exercise 4.5.11 Prove the Ergodic Theorem for a finite state space Markov chain; that is the result that for an irreducible
Markov chain {xt } living in a finite space X, which has a unique invariant probability measure µ, the following applies
almost surely:
T
1X X
lim f (xt ) = f (i)µ(i),
T →∞ T
t=1 i

for every f : X → R.
Hint: You may proceed as follows. Define a sequence of empirical occupation measures for T ∈ N, A ∈ B(X):
T −1
1 X
vT (A) = 1{Xt ∈A} , ∀A ∈ B(X).
T t=0

Now, define:
t
X X 
Ft (A) = 1{Xs ∈A} − t P (A|x)vt (x)
s=1 X
80 4 Martingales, Foster-Lyapunov Criteria for Stability of Markov Chains, and Stochastic Iterative Dynamics
t
X t−1 X
X 
= 1{Xs ∈A} − P (A|x)1{Xs =x} (4.38)
s=1 s=0 X

Let Ft = σ(X0 , · · · , Xt ). Verify that, for t ≥ 2,

E[Ft (A)|Ft−1 ]
 X t t−1 X
X  
=E 1{Xs ∈A} − P (A|x)1{Xs =x} Ft−1
s=1 s=0 X
 X  
= E 1{Xt ∈A} − P (A|x)1{Xt−1 =x} Ft−1
X
t−1
X t−2 X
X 
+ 1{Xs ∈A} − P (A|x)1{Xs =x}
s=1 s=0 X
t−1
X t−2 X
X  
=0+ 1{Xs ∈A} − P (A|x)1{Xs =x} Ft−1 (4.39)
s=1 s=0 X
= Ft−1 (A), (4.40)

where the last equality follows from the fact that E[1{Xt ∈A} |Ft−1 ] = P (Xt ∈ A|Ft−1 ). Furthermore,

|Ft (A) − Ft−1 (A)| ≤ 1.

Now, we have a sequence which is a martingale sequence. We will invoke a martingale convergence theorem; which is
applicable for martingales with bounded increments. By a version of the martingale stability theorem, it follows that
1
lim Ft (A) = 0.
t→∞ t
You need to now complete the remaining steps.
Hint: You can use the Azuma-Hoeffding inequality (Theorem 4.1.15) [97] and the Borel-Cantelli Lemma to complete the
steps.
We note that a similar argument could also be made for countably infinite X or uncountable X under additional conditions.

Exercise 4.5.12 Let τ be a stopping time with respect to the filtration Ft . Let Xn be a (discrete-time) sequence of random
variables so that each Xn is Fn -measurable. Show that Xτ is Fτ -measurable.
Hint: We need to show that for every real a: {Xτ ≤ a} ∩ {τ ≤ k} ∈ Fk . Observe that {Xτ ≤ a} ∩ {τ ≤ k} =
∪km=0 {Xτ ≤ a} ∩ {τ = m} and that for each m, {Xτ ≤ a} ∩ {τ = m} ∈ Fm ⊂ Fk .

Exercise 4.5.13 To appreciate that the condition of measurability on control or estimation policies is not a superfluous
one, read the paper [324]: G. L. Wise. A note on a common misconception in estimation. Systems & Control letters, 1985:
355-356.
5

Optimal Stochastic Control with Finite and Discounted Infinite Horizons and
Dynamic Programming

In this chapter, we introduce the method of dynamic programming for controlled stochastic systems, and consider optimal
stochastic control problems under finite horizon and discounted infinite horizon expected cost criteria.
Recall that a fully observed Markov control model is a five-tuple

(X, U, {U(x), x ∈ X}, T , c)

such that X is the (standard Borel) state space, U is the action space, U(x) ⊂ U is the control action set when the state is
x, so that
K = {(x, u) : x ∈ X, u ∈ U(x)} ⊂ X × U,
is the set of feasible state-action pairs. T is a stochastic kernel on X given K. Finally c : K → R is the cost function.
One also can have a dependence of the cost function on the time variable so that ct can be the cost at time t or the action
set U(x) can also depend on time. In this case one can add the time variable t, as a further component, to the state variable
x, with a deterministic evolution for the time variable. Conceptually, such a generalization does not introduce any further
obstacles for finite horizon problems. Often, ct ≡ c, that is c does not depend on time (however, there may be a terminal
cost different from c, to be considered).
Let, as in Section 2.2.1, ΓA denote the set of all admissible policies. Let γ = {γt , 0 ≤ t ≤ N − 1} ∈ ΓA be a policy.
Consider the following expected cost:
N
X −1
JN (x, γ) := Exγ [ c(xt , ut ) + cN (xN )], (5.1)
t=0

where cN (.) is the terminal cost function. Define



JN (x) := inf JN (x, γ)
γ∈ΓA

5.1 Dynamic Programming, Optimality of Markov Policies and Bellman’s Principle of


Optimality


The goal is to find, if there exists one, an admissible policy such that JN (x) is attained; this will be an optimal policy.
We note that the infimum value, in general, may not be attained by some policy. In the following, we will also present
conditions which will ensure the existence of optimal policies.
82 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

5.1.1 Backwards Induction

Let ht = {x[0,t] , u[0,t−1] } denote the history process for t ∈ N. By Theorem 4.1.3, provided that the cost is integrable
under the induced probability measure given a policy, we note that the cost can be expressed as:

JN (x, γ) = Exγ c(x0 , u0 )

+E γ c(x1 , u1 )

+E γ c(x2 , u2 )
+...
  
γ
+E [c(xN −1 , uN −1 ) + cN (xN )|hN −1 ] hN −2 . . . h1 x0 = x ,

γ0 ,··· ,γN −1
= Ex c(x0 , u0 )

γ1 ,··· ,γN −1
+E c(x1 , u1 )

γ2 ,··· ,γN −1
+E c(x2 , u2 )
+...
  
+E γN −1 [c(xN −1 , uN −1 ) + cN (xN )|hN −1 ] hN −2 . . . h1 x0 = x ,

= Exγ0 c(x0 , u0 )

+E γ1 c(x1 , u1 )

+E γ2 c(x2 , u2 )
+...
  
γN −1
+E [c(xN −1 , uN −1 ) + cN (xN )|hN −1 ] hN −2 . . . h1 x0 = x ,

Thus, by the equalities above, we obtain:



inf JN (x, γ) = inf Exγ0 c(x0 , u0 )
γ∈ΓA γ0

+ inf E γ1 c(x1 , u1 )
γ1

+ inf E γ2 c(x2 , u2 )
γ2
+...
  
γN −1
+ inf E [c(xN −1 , uN −1 ) + cN (xN )|hN −1 ] hN −2 . . . h1 x0 = x , (5.2)
γN −1

The discussion above reveals that we can start with the final time stage, obtain a solution for γN −1 and move backwards
for t ≤ N − 2. To this end, we present a critical supporting result in the following.
5.1 Dynamic Programming, Optimality of Markov Policies and Bellman’s Principle of Optimality 83

5.1.2 Optimality of Deterministic Markov Policies

We will observe that when there is an optimal solution, the optimal solution can be taken to be Markov. Even when an
optimal policy may not exist, any measurable policy can be replaced with one which is Markov, under fairly general
conditions, as we discuss below. In the following, first, we will follow David Blackwell’s [48] and Hans Witsenhausen’s
[330] ideas to obtain a very interesting result.

Theorem 5.1.1 (Blackwell’s Theorem on Redundancy of Information beyond the State) Let X, Y, U be complete, sep-
arable, metric spaces, and let P be a probability measure on B(X × Y), and let c : X × U → R be a Borel measurable and
bounded cost function. Then, for any Borel measurable function γ : X × Y → U, there exists another Borel measurable
function γ ∗ : X → U such that
Z Z
c(x, γ ∗ (x))PX (dx) ≤ c(x, γ(x, y))P (dx, dy)
X X×Y

where PX is the marginal of P on X. Thus, policies based only on x almost surely, are optimal.

Proof. We will construct a γ ∗ given γ. Let u = γ(x, y). To emphasize the random nature of the variables considered, let
us again denote with capital letters X, Y, U the random variables whose realizations are x, y and u, respectively. Given γ,
we write for any Borel D ⊂ U and x ∈ X,
Z
P γ (U ∈ D|x) = P (γ(X, Y ) ∈ D|X = x) = 1{γ(x,y)∈D} P (Y ∈ dy|X = x).
Y

We then have Z Z Z 
γ
c(x, γ(x, y))P (dx, dy) = c(x, u)P (du|x) P (dx),
X×Y X U

Consider
Z
γ
h (x) := c(x, u)P γ (du|x) (5.3)
U

• Suppose the space U is countable. In this case, let us enumerate the elements in U as {ui , i = 1, 2, . . . }. Then, we
could define:
Di = {x ∈ X : c(x, ui ) ≤ hγ (x)}, i = 1, 2, . . . .
We note that X = i Di : Suppose not, then ∃x ∈ X with c(x, ui ) > hγ (x) for all i ∈ N, and thus for this x:
S

X 
hγ (x) = c(x, u)P γ (du|x) > hγ (x), (5.4)
U

leading to a contradiction. Now define,

γ ∗ (x) = uk if x ∈ Dk \ (∪k−1
i=1 Di ), k = 1, 2, . . . ,

Such a function is measurable, by construction and performs at least as good as γ.


• We now provide a proof for the actual statement. Let D = {(x, u) ∈ X × U : c(x, u) ≤ hγ (x)}. D is a Borel set
since c(x, u) − hγ (x) is Borel. Define Dx = {u ∈ U : (x, u) ∈ D} for all x ∈ X. Now, for every element x we can
pick a member u which is in D; this Rdefines a map from X to U. The question now is whether the constructed map is
Borel measurable. Now, for every x, 1{γ(x,y)∈D} P (dy|x) > 0 by the relation (5.3), since otherwise we would arrive
at a contradiction via (5.4). Then, by a measurable selection theorem of Blackwell and Ryll-Nardzewski [51] (see also
p. 255 of [116]), there exists a Borel-measurable function γ ∗ : X → U such that its graph is contained in D, that is,
{(x, γ ∗ (x)) ∈ D}.


84 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

The following culminate into the following critical result.


PN −1
Theorem 5.1.2 Let {(xt , ut )} be a controlled Markov chain. Consider (5.1), that is, the minimization of E γ [ t=0 c(xt , ut )+
cN (xN )], over all admissible control policies. Any such policy can be replaced with one which is (deterministic) Markov
and which is at least as good as the original policy. In particular, if an optimal control policy exists, there is no loss in
restricting policies to be Markov.

Proof. The proof follows from a sequential application of Theorem 5.1.1, starting with the final time stage. For any admis-
sible policy, the cost Z
E[c(xN −1 , γN −1 (hN −1 )) + cN (z)T (dz|xN −1 , γN −1 (hN −1 ))],
X

can be replaced with a measurable policy γN −1
Z
∗ ∗
E[c(xN −1 , γN −1 (xN −1 )) + cN (z)T (dz|xN −1 , γN −1 (xN −1 ))],
X

which leads to a cost that is at least as good as one obtained with γN −1 .


We define  Z 
∗ ∗
JN −1 (xN −1 ) := E c(xN −1 , γN −1 (xN −1 )) + cN (z)T (dz|xN −1 , γN −1 (xN −1 )) ,
X

and consider then, in view of (5.2),


 Z 
E c(xN −2 , γN −2 (hN −2 )) + JN −1 (z)T (dz|xN −2 , γN −2 (hN −2 )) .
X


This expression can also be lower bounded by a measurable Markov policy γN −2 so that the expected cost
 Z 
∗ ∗
JN −2 (xN −2 ) := E c(xN −2 , γN (x
−2 N −2 )) + J N −1 (z)T (dz|x , γ (x
N −2 N −2 N −2 )) .
X

is lower than that achieved by the admissible policy. By induction, for all time stages, one can replace the policies with
a deterministic Markov policy which leads to a cost which is at least as desirable as the cost achieved by the admissible
policy. ⋄
Given this result, we have the following important optimality principle.

5.1.3 Bellman’s principle of optimality and Dynamic Programming

Consider 5.1. Let {Jt (xt )} be a sequence of functions on X defined by

JN (x) = cN (x)

and for 0 ≤ t ≤ N − 1 Z
Jt (x) = min {c(x, u) + Jt+1 (z)T (dz|x, u)}.
u∈Ut (x) X

Let there be minimizing measurable functions which are deterministic, denoted by {ft (x)}, so that
Z
Jt (x) = c(x, ft (xt )) + Jt+1 (z)T (dz|x, ft (x))}
X

Then we have the following:

Theorem 5.1.3 The policy γ ∗ = {f0 , f1 , . . . , fN −1 } is optimal and the optimal expected cost function (also called the
value function) is equal to
5.1 Dynamic Programming, Optimality of Markov Policies and Bellman’s Principle of Optimality 85

JN (x) = J0 (x)

Proof. We compare the expected cost generated by the above policy, with respect to the cost obtained by any other policy,
which can be taken to be deterministic Markov in view of Theorem 5.1.2.
We provide the proof by a backwards induction method in view of (5.2). Consider the time stage t = N − 1. For this stage,
the optimal cost (or, value) is equal to
Z
JN −1 (x) = min{c(x, u) + cN (z)T (dz|xN −1 = x, uN −1 = u)}
u X


Suppose there is a cost CN −1 (x), achieved by some policy η = {ηk , k ∈ {0, 1, · · · , N − 1}}, which we take to be
deterministic Markov (without loss). Since,

CN −1 (x)
Z
= c(x, ηN −1 (x)) + cN (z)T (dz|xN −1 = x, uN −1 = ηN −1 (xN −1 ))
X
≥ JN −1 (x)
Z
= min{c(x, u) + cN (z)T (dz|xN −1 = x, uN −1 = u)}, (5.5)
u X


it must be that CN −1 (x) ≥ JN −1 (x). Now, we move to time stage N − 2. In this case, the cost is given by
Z
∗ ∗
CN −2 (x) = c(x, η(xN −2 )) + CN −1 (z)T (dz|xN −2 = x, uN −2 = η(xN −2 ))
X
Z
≥ min{c(x, u) + JN −1 (z)T (dz|xN −2 = x, uN −2 = u)}
u X
=: JN −2 (x)

where the inequality is due to the fact that CN −1 (x) ≥ JN −1 (x) and the minimization. We can, by induction, show that
the recursion holds for all 0 ≤ t ≤ N − 2.

5.1.4 Examples

Example 5.1 (Dynamic Programming and Investment). [164, Section 3.6] A investor’s wealth dynamics is given by the
following:
xt+1 = ut wt ,
where {wt } is an i.i.d. R+ -valued stochastic process with E[wt ] = w̄. The investor has access to the past and current
wealth information and his actions. The goal is to maximize, for some b > 0,
T
X −1
J(x0 , γ) = Exγ0 [ b(xt − ut )].
t=0

The investor’s action set for any given x is: U(x) = [0, x]. We will find an optimal admissible policy.
For this problem, the state space is R+ , the control action space at state x is [0, x], the information at the controller is
It = {x[0,t] , u[0,t−1] }. The kernel is described by the relation xt+1 = ut wt . Using Dynamic Programming

JT −1 (x) = max E[b(xT −1 − uT −1 )|xT −1 = x, uT −1 = u]


u∈[0,xT −1 ]

= max b(x − u) = b(x). (5.6)


u∈[0,x]
86 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

Since there is no more future, the investor needs to collect the wealth at time T − 1, that is uT −1 = 0. For t = T − 2

JT −2 (x) = max E[b(x − u) + JT −1 (xT −1 )|xt−2 = x, uT −2 = u]


u∈[0,x]

= max E[b(x − u) + bxT −1 |xt−2 = x, uT −2 = u]


u∈[0,x]
 
= max b(x − u) + bE[wT −2 ]u
u∈[0,x]
 
= max bx + b(w̄ − 1)u
u∈[0,x]

It follows then that if w̄ > 1, uT −2 = xT −2 (that is, investment is favourable), otherwise uT −2 = 0. Recursively, one
concludes that if w̄ > 1, ut = xt is optimal until t = T − 1, at t = T − 1, uT −1 = 0, leading to J0 (x0 ) = bw̄T −1 x0 .
If w̄ < 1, it is optimal to collect at time 0, that is u0 = 0, leading to J0 (x0 ) = bx0 . If w̄ = 1, both of these policies lead to
the same reward.

Example 5.2 (Linear Quadratic Systems). Consider the following Linear Quadratic (LQ) problem with q > 0, r > 0, pT >
0:
T
X −1
inf Exγ [ qx2t + rt2 + pT x2T ]
γ
t=0

for a linear system:


xt+1 = axt + ut + wt ,
2
where wt is a zero-mean random variable with variance σw < ∞. We can show, by the method of completing the squares,
that:
T
X −1
Jt (xt ) = Pt x2t + 2
Pt+1 σw
k=t

where
2
Pt+1 a2
Pt = q + Pt+1 a2 −
Pt+1 + r
and the optimal control policy is
−Pt+1 a
ut = xt .
Pt+1 + r
Note that, the optimal control policy is Markov (as it uses only the current state). For a more general treatment for such LQ
problems, see Section 5.3. A typical setup is the case where wt is Gaussian; in this case the problem above is often referred
to as the Linear Quadratic Gaussian (LQG) optimal control problem.

5.2 Existence of Minimizing Selectors and Measurability

The above dynamic programming arguments hold when there exist minimizing control policies (selectors measurable with
respect to the Borel σ-field on X). The following results build on [168, Theorem 2], [283], [282] and [201] (see Appendix
C). We also refer the reader to [164] for a comprehensive analysis and detailed literature review and [125, Theorem 2.1].
Measurable Selection Hypothesis: Given a sequence of functions Jt : X → R, there exists
Z
Jt (x) = min (c(xt , ut ) + Jt+1 (y)T (dy|x, u)),
ut ∈Ut (x) X

for all x ∈ X, for t ∈ {0, 1, 2 . . . , N − 1} with

JN (xN ) = cN (xN ).
5.2 Existence of Minimizing Selectors and Measurability 87

Furthermore, there exist measurable functions ft such that


Z
Jt (x) = c(xt , ft (xt )) + Jt+1 (y)T (dy|x, ft (xt )),
X


Recall that a set in a normed linear space is (sequentially) compact if every sequence in the set has a converging subse-
quence.

R
Assumption 5.2.1 (Condition WF) (i) For every continuous and bounded v on X (that is, v ∈ Cb (X)), X
T (dy|x, u)v(y)
is a continuous function on X × U (in this case, we call T a weakly continuous transition kernel).
(ii) The cost function to be minimized c(x, u) is bounded and continuous on both U and X.
(iii)If applicable, cN is continuous and bounded.
(iv) Ut (x) = U is compact.

R
Assumption 5.2.2 (Condition S) (i) For every measurable and bounded v on X (that is, v ∈ L∞ (X; R)), X T (dy|x, u)v(y)
is a continuous function on U, for every fixed x (in this case, we call T a strongly continuous transition kernel in u for
every fixed x).
(ii) For every x ∈ X the bounded measurable cost function c(x, u) is continuous on U,
(iii)If applicable, cN is bounded measurable.
(iv) Ut (x) = U is compact.

Theorem 5.2.1 Under Assumption 5.2.1 or Assumption 5.2.2, there exists an optimal solution and the measurable selection
hypothesis applies, and there exists a minimizing control policy ft : X → U. Furthermore, under Assumption 5.2.1, Jt is
continuous for any t ≥ 0.

The result follows from the following three lemmas below:

Lemma 5.2.1 A continuous function f : X → R over a compact set A ⊂ X admits a minimum.

Proof. Let δ = inf x∈A f (x). Let {xi } be a sequence such that f (xi ) converges to δ. Since A is compact {xi } must have
a converging subsequence {xi(n) }. Let the limit of this subsequence be x̄. Then, it follows that, {xi(n) } → x̄ and thus, by
continuity {f (xi(n) )} → f (x̄). As such f (x̄) = δ. ⋄
1
To see why compactness is important, consider inf x∈A x for A = [1, 2) or A = R. In both cases there does not exist an x
value in the specified set which attains the infimum.

Lemma 5.2.2 Let U be compact, and c(x, u) be continuous on X × U. Then, minu∈U c(x, u) is continuous on X.

Proof. Let xn → x, un optimal for xn and u optimal for x. Such optimal action values exist as a result of compactness of
U and continuity. Now,

| min c(xn , u) − min c(x, u)|


u u
 
≤ max c(xn , u) − c(x, u), c(x, un ) − c(xn , un ) (5.7)

The first term above converges to zero since c is continuous in x, u. The second converges also. Suppose otherwise. Then,
for some ϵ > 0, there exists a subsequence such that
88 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

|c(x, ukn ) − c(xkn , ukn )| ≥ ϵ

Consider the sequence (xkn , ukn ). There exists a further subsequence (in this sequence (xkn , ukn )) (xkn′ , ukn′ ) which
converges to x, u′ for some u′ since U is compact. Hence, for this subsequence, we have convergence of c(xkn′ , ukn′ ) as
well as c(x, ukn′ ) to the same term, leading to a contradiction. ⋄

Lemma 5.2.3 Let c(x, u) be a continuous function on U for every x, where U is a compact set. Then, there exists a Borel
measurable function f : X → U such that
c(x, f (x)) = min c(x, u)
u∈U

Proof. A sketch is as follows: Let c̃(x) := minu∈U c(x, u). The function

c̃(x) := min c(x, u),


u∈U

is Borel measurable. This follows from the observation that it is sufficient to prove that {x : c̃(x) > α} is Borel for every
α ∈ R. By continuity of c and compactness of U, with a successively refining quantization of the space of control actions U
(such a sequence of quantizers map U to a sequence of finite sets (expanding as n increases), so that limn→∞ supu |Qn (u)−
u| = 0 and the cardinality |Qn (U)| < ∞ for every n)
\ \
{x : c̃(x) ≥ α} = {x : c(x, Qn (u)) ≥ α}
n Qn (u),u∈U

the result follows since each of {x : c(x, Qn (u)) ≥ α} is Borel. Define F := {(x, u) : c(x, u) = c̃(x), x ∈ X.} This set
is a Borel set and for every x, {u : (x, u) ∈ F} is a closed set. The question is now whether one can construct a measurable
(selection) function γ in F so that {(x, γ(x)), x ∈ X} ⊂ F. One can construct a measurable function which lives in this
set, using the property that U is a separable metric space: This builds on measurable selection results, e.g. Schäl [283]
and [201]; see Theorem C.0.1 (building on [168, Theorem 2], [283], [282] and [201], among others; see Appendix C).

5.2.1 Some Relaxations on the Measurable Selection Conditions

We first note that one can replace the compactness condition with an inf-compactness condition, and modify Condition 1
in Assumption 5.2.1 as below:

Assumption 5.2.3 (Condition 3) For every x ∈ X the cost function to be Rminimized c(x, u) is continuous on X × U; is
non-negative; {u : c(x, u) ≤ α} is compact for all α > 0 and all x ∈ X; X T (dy|x, u)v(y) is a continuous function on
X × U for every continuous and bounded v.

Theorem 5.2.2 Under Assumption 5.2.3, the Measurable Selection Hypothesis applies.

The measurable selection results also hold when U(x) depends on x so that it is compact for each x and {(x, u) : u ∈
U(x), x ∈ X} is a Borel subset of X × U:

Lemma 5.2.4 [168, Theorem 2], [283] [201] Let X, U be standard Borel spaces and Υ = (x, ψ(x)) where ψ(x) ⊂ U be
such that, ψ(x) is compact for each x ∈ X and Υ is a Borel measurable set in X × U. Let c(x, u) be a continuous function
on ψ(x) for every x.

(i) Then, there exists a Borel measurable function f : X → U such that

c(x, f (x)) = min c(x, u)


u∈ψ(x)
5.3 The Linear Quadratic Regulator (LQR) Problem 89

(ii) If continuity is also to be attained for the value function c(x, f (x)) (a close look at the proof of Lemma 5.2.2 reveals
that) it suffices if U(x) is compact and U(x) is an upper semi-continuous set-valued function (the implication being
that: for any xn → x and u′n ∈ U(xn ), there exists a subsequence u′nk which converges to some u′ with the property
that u′ ∈ U(x)) and c is continuous.

We could relax the continuity condition and change it with lower semi-continuity. A function is lower semi-continuous
at x0 if lim inf x→x0 f (x) ≥ f (x0 ). We state the following, see also [164, Theorem 3.3.5] (we note there is a slight typo
in [164, Theorem 3.3.5]; in [164, Condition 3.3.2.(c2)] should be assumed and only [164, Condition 3.3.2(c1)] is not
sufficient for [164, Condition 3.3.2] to imply measurable selection).

Theorem 5.2.3 The following hold:

(a) Suppose that (i) U(x) is compact for every x and {(x,
R u) : u ∈ U(x)} is a Borel subset of X × U, (ii) c is lower
semi-continuous on U(x) for every x ∈ X, and (iii) v(xt+1 )P (dxt+1 |xt = x, ut = u) is lower semi-continuous on
U(x) for every x ∈ X and every measurable and bounded v on X. Then, the measurable selection hypothesis applies.

R c is lower semi-continuous on {(x, u) : u ∈ U(x), x ∈ X}, (ii) for every lower semi-continuous function v on
(b) If (i)
X, v(xt+1 )P (dxt+1 |xt = x, ut = u) is lower semi-continuous on {(x, u) : u ∈ U(x), x ∈ X}, and (iii) U(x) is
compact for every x ∈ X and U(x) is an upper semi-continuous set-valued function; then the value function v is lower
semi-continuous.

For further related relaxations, see Appendix C and [164, Appendix D].
Universally Measurable Policies. As we discuss in Appendix C, studying the class of universally measurable and semi-
analytic functions allows one to even further relax conditions required for carrying out dynamic programming recursions
(and integrations) with regard to their well-posedness properties and for arriving at ϵ-optimal policies via dynamic pro-
gramming.
For many problems, one can compute an optimal solution directly, without explicitly studying existence. The linear
quadratic setup is one such important case.

5.3 The Linear Quadratic Regulator (LQR) Problem

Consider the following linear system

xt+1 = Axt + But + wt , (5.8)

where x ∈ Rn , u ∈ Rm and w ∈ Rn . Suppose {wt } is i.i.d. zero-mean with a given covariance matrix E[wt wtT ] = W for
all t ≥ 0 (not necessarily Gaussian).
The goal is to obtain
inf J(x, γ),
γ∈ΓA

where
N
X −1
J(x, γ) = Exγ [ xTt Qxt + uTt Rut + xTN QN xN ], (5.9)
t=0

with R = RT > 0, Q = QT ≥ 0, QN = QTN ≥ 0 (where, for matrices, the notations > and ≥ denote the positive-definite
and positive semi-definite properties, respectively).

Theorem 5.3.1 Consider (5.9). The optimal control is linear and has the form:
90 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

ut = −(B T Pt+1 B + R)−1 B T Pt+1 Axt

where Pt solves the Discrete-Time Riccati Equation:

Pt = Q + AT Pt+1 A − AT Pt+1 B(B T Pt+1 B + R)−1 B T Pt+1 A, (5.10)

with final condition PN = QN . The optimal cost is given by


N
X −1
J(x0 ) = xT0 P0 x0 + E[wtT Pt+1 wt ]
t=0

In the following, we study the Riccati equation (5.10). Consider the linear system

xt+1 = Axt + But , yt = Cxt (5.11)

Here, yt is a measurement variable and xt is an Rn -valued state variable. Such a system is said to be controllable [82],
if for any initial xi and a final xf , there exists T ∈ N and a sequence of control actions u0 , u1 , · · · , uT −1 such that with
x0 = xi , we have xT = xf . If xf is restricted to be 0 ∈ Rn , and the above holds (but possibly with T → ∞), the system
is said to be stabilizable. Thus, the only modes in a stabilizable system that are not controllable are the stable ones.
Now let B = 0 in (5.11). Such a system is said to be observable if by measuring y0 , y1 , · · · , yT , for some T ∈ N, x0 can
be uniquely recovered. Such a system is called detectable if all unstable modes of A are observable, in the sense that if
yt → 0, it must be that xt → 0.
There are well-known algebraic tests to verify controllability and observability. A very useful result building on the Cayley-
Hamilton theorem is that if a system cannot be moved from any initial state to any final state in n (that is, the dimension
of Rn ) time stages, the system is not controllable; and if a system’s initial state cannot be recovered by having the n
measurements {y0 , y1 , · · · , yn−1 }, the system is not observable. In particular, the linear system above with matrices (A, B)
is controllable if and only if
B AB · · · An−1 B
 

is full-rank. The pair (A, C) is observable if and only if (AT , C T ) is controllable.


For a review of linear systems theory, the reader is referred to, e.g. [82].

Theorem 5.3.2 (i) If (A, B) is controllable there exists a solution to the Riccati equation

P = Q + AT P A − AT P B(B T P B + R)−1 B T P A.

(ii) if (A, B) is controllable and, with Q = C T C, (A, C) is observable; as t → −∞ (or as N → ∞ with QN = P̄ fixed
for an arbitrary positive semi-definite matrix P̄ ), the sequence of Riccati recursions,

Pt = Q + AT Pt+1 A − AT Pt+1 B(B T Pt+1 B + R)−1 B T Pt+1 A,

converges to some limit P that satisfies

P = Q + AT P A − AT P B(B T P B + R)−1 B T P A.

That is, convergence takes place for any initial condition P̄ . Furthermore, such a P is unique, and is positive definite.
Finally, under the optimal stationary control policy

ut = −(B T P B + R)−1 B T P Axt ,

the solution to xt+1 = Axt + But is stable; i.e., xt → 0.


(iii)Under the conditions of part (ii), the stationary policy above minimizes,
5.3 The Linear Quadratic Regulator (LQR) Problem 91
N −1
1 γ X T
lim sup Ex [ xt Qxt + uTt Rut ], (5.12)
N →∞ N t=0

for the system (5.8), for every x ∈ Rn . Furthermore, the optimal cost is E[wT P w] = Trace(P W ).

Remark 5.3. Part (i) can be relaxed to (A, B) being stabilizable; and part (ii) to (A, C) being detectable for the existence
of a unique P and a stable system under the optimal policy. In this case, however, P may only be positive semi-definite.

Proof.

(i) Assume that wt = 0 for all t; the noise does not affect the recursions in the Riccati equation. Now, since the system is
controllable there exists a control sequence such that xt = 0 for t ≥ n which also satisfies ut = 0 for t ≥ n. The cost
P ∞ T T
t=0 xt Qxt + ut Rut induced by this control sequence is finite (and thus bounded by some M (x0 )). Now, define
(N )
P0 through
N −1
(N )
X
xT0 P0 x0 = inf Exγ [ xTt Qxt + uTt Rut ]
γ∈ΓA
t=0

and observe that


(N ) (N +1)
xT0 P0 x0 ≤ xT0 P0 x0 ≤ M (x0 ).
(N )
As a result, for a fixed x0 , we can conclude that the sequence {xT0 P0 x0 , T ≥ 0} is monotone (non-decreasing) and
 T
bounded from above. Thus, the sequence has a limit. By selecting different values of x0 (e.g., with x0 = 1 0 0 · · · 0 ,
 T  T (N )
x0 = 0 1 0 · · · 0 , x0 = 1 1 0 · · · 0 and so on), we conclude that there is a fixed point P such that xT0 P0 x0 →
(N )
xT0 P x0 for any x0 ∈ Rn (i.e., P0 → P point-wise in the matrix entries).
(ii) As above, assume again that wt = 0 for all t ≥ 0. Let P be the fixed point in (i). We will show that this is the unique
fixed point.
We use the property that, through a change of limit supremum and infimum argument (as in Lemma 5.5.1 further
below),
N
X −1
∞ > M (x) ≥ inf lim sup Exγ [ xTt Qxt + uTt Rut ]
γ∈ΓA N →∞
t=0
N
X −1
≥ lim sup inf Exγ [ xTt Qxt + uTt Rut ] = xT P x (5.13)
N →∞ γ∈ΓA t=0

Now, we show that, the sequence of policies that are optimal for each N converge to the stationary policy by γ ∗ (xt ) =
−(B T P B + R)−1 B T P Axt and that this policy attains the cost xT P x: Note that, with x0 = x,
N −1
∗ X
xT P x = E γ [( xTt Qxt + uTt Rut ) + xTN P xN ] (5.14)
t=0
P∞
is finite and as a result, under this policy γ ∗ , we have that t=0 xTt Qxt + uTt Rut is finite.
However, since the induced cost is finite and R > 0, under this control policy, ut → 0. Therefore, the policy γ ∗

ut = −(B T P B + R)−1 B T P Axt = γ ∗ (xt ),

is stabilizing: This follows because since xT Qxt → 0 (and ut → 0), by observability of (C, A) it must be that xt → 0
as well (note that here one should also use that ut → 0). As a result, we conclude that, by taking N → ∞ in (5.14), γ ∗
satisfies

γ∗
X
T
x Px = E [ xTt Qxt + uTt Rut ],
t=0
92 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

and is therefore optimal (by 5.13).


Given this optimality (which will be useful also for (iii) below), we now show uniqueness. Let
NX−1
xT0 P0N,P̄ x0 := inf γ
Ex [ xTt Qxt + uTt Rut + xTN P̄ xN ] (5.15)
γ∈ΓA
t=0

be the solution of the optimization problem where PN = P̄ . We will show that P0N,P̄ → P regardless of the value of
the positive semi-definite matrix P̄ , leading to the uniqueness of the limit.
∗ ∗ ∗
By writing Exγ0 [xTN P̄ xN ] = Exγ0 [xTN P xN ] + Exγ0 [xTN (P̄ − P )xN ], noting that, as P is a solution to the Riccati
recursion,
N −1
∗ X
Exγ [ xTt Qxt + uTt Rut + xTN P xN ] = xT0 P x0 ,
t=0

we have that
N −1
(N ) ∗ X ∗
xT0 P0 x0 ≤ xT0 P0N,P̄ x0 ≤ Exγ [ xTt Qxt + uTt Rut + xTN P xN ] + Exγ0 [xTN (P̄ − P )xN ].
t=0

or ∗
(N )
xT0 P0 x0 ≤ xT0 P0N,P̄ x0 ≤ xT0 P x0 + Exγ0 [xTN (P̄ − P )xN ].
The above holds as γ ∗ is not necessarily optimal, and provides an upper bound, for (5.15). However, through the
property that xN → 0 as N → ∞ under ut = γ ∗ (xt ), we conclude that P0N,P̄ → P and hence uniqueness follows.
(iii)As in (ii), we use the property that, through a change of limit supremum and infimum argument (as in Lemma 5.5.1
further below),
N −1
1 γ X T
inf lim sup Ex [ xt Qxt + uTt Rut ]
γ∈ΓA N →∞ N t=0
N −1
1 X
≥ lim sup inf Exγ [ xTt Qxt + uTt Rut ] (5.16)
N →∞ N γ∈ΓA t=0

PN −1
Since, inf γ∈ΓA Exγ [ t=0 xTt Qxt + uTt Rut ] is determined by P (N ) that converges to P , leading to the optimality of
γ ∗ , and the policy ut = −(BP B + R)−1 B T P Axt is stabilizing, this implies that the policy γ ∗ is optimal for (5.12) as
well; see (5.13) and the following discussion. The optimal cost then is E[wT P w] = Trace(P W ), via observing that
(N )
Pt+1 → P for all t as N → ∞ and writing
 N −1 
1 (N )
X (N )
lim xT0 P0 x0 + E[wtT Pt+1 wt ] = E[wT P w].
N →∞ N
t=0


We will discuss average cost optimization problems in further detail in Chapter 7.

5.4 Optional: A Strategic Measures Approach

For stochastic control problems, strategic measures are defined (see [283], [116] and [123]) as the set of probability
measures induced on the product spaces of the state and action pairs by measurable control policies: Given an initial
distribution on the state, and a policy, one can uniquely define a probability measure on the product space. Topological
5.4 Optional: A Strategic Measures Approach 93

properties, such as measurability and compactness, of sets of strategic measures are studied in [283], [116], [123]
and [49].
We assume, as before, that the spaces considered are standard Borel. In the following, we consider a finite horizon
problem, with time horizon N − 1.

Theorem 5.4.1 Let LR (µ) be the set of strategic measures induced by (possibly randomized) ΓA with x0 ∼ µ. Then,
for any P ∈ LR (µ), there exists an augmented space Ω and a probability measure η such that
Z
P (B) = η(dω)Pµγ(ω) (B), B ∈ B((X × U)N ),

where each γ(ω) ∈ ΓA is deterministic admissible.

Proof. Here, we build on Lemma 1.2 in Gikhman and Shorodhod [143] and Theorem 1 in [122]. Any stochastic
kernel P (dx|y) can be realized by some measurable function x = f (y, v) where v is a uniformly distributed random
variable on [0, 1] and f is measurable (see also [56] for a related argument). One can define a new random variable
(ω = (v0 , v1 , · · · , vT −1 )). In particular, η can be taken to be the probability measure constructed on the product space
[0, 1]N by the independent variables vk , k ∈ {0, 1, · · · , N − 1}. ⋄
One implication of this theorem is that if one relaxes the measure η to be arbitrary, a convex representation would be
possible. That is, the set
Z
P (B) = η(dω)Pµγ(ω) (B), B ∈ B((X × U)N ), η ∈ P(Ω)

is convex, when one does not restrict η to be a fixed measure. Furthermore, the extreme points of these convex sets
consist of policies which are deterministic. A further implication then is that, since the expected cost function is linear
in the strategic measures, one can without any loss consider the extreme points while searching for optimal policies. In
particular,
inf J(x, γ) = inf J(x, γ)
γ∈ΓM R γ∈ΓM

and
inf J(x, γ) = inf J(x, γ).
γ∈ΓAR γ∈ΓA

Thus, deterministic policies are as good as any other. This is not surprising in view of Theorem 5.1.1.
We present the following characterization for strategic measures. Let for all n ∈ N, hn = {x0 , u0 , · · · , xn−1 , un−1 , xn , un },
and P (dxn |hn−1 ) = T (dxn |xn−1 , un−1 ) be the transition kernel.
Let LA (µ) be the set of strategic measures induced by deterministic policies and let LR (µ) be the set of strategic
measures induced by independently provided randomized policies. Such an individual randomized policy can be rep-
resented in a functional form, as noted earlier: for any stochastic kernel Π k from Yk to Uk , there exists a measurable
function γ k : [0, 1] × Yk → Uk such that

m{r : γ k (r, y k ) ∈ A} = γ k (uk ∈ A|y k ), (5.17)

and m is the uniform distribution (Lebesgue measure) on [0, 1].


 
QN
Theorem 5.4.2 A probability measure P ∈ P k=1 (X × U) is a strategic measure induced by a randomized policy

(that is in LR (µ)) if and only if for every n ∈ N and for all continuous and bounded g:
Z Z Z 
P (dhn−1 , dxn )g(hn−1 , xn ) = P (dhn−1 ) g(hn−1 , z)T (dz|hn−1 ) ,
X
(5.18)
94 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming
R
(where we recall that T (B|hn−1 ) = T (dxn |xn−1 , un−1 )), and
B
Z Z Z 
P (dhn )g(hn−1 , xn , un ) = P (dhn−1 , dxn ) g(hn−1 , xn , a)γ n (da|hn−1 , xn ) ,
Un
(5.19)

for some stochastic kernel γ n on Un given hn , xn , with P (dω0 ) = µ(dw0 ).

Proof. The proof follows from the fact that testing the equalities such as (5.18-5.19) on continuous and bounded
functions implies this property for any measurable and bounded function (that is, continuous and bounded functions
form a separating class, see e.g. [42, p. 13] or [119, Theorem 3.4.5]) ⋄
An implication is the following.

Theorem 5.4.3 [283] The set of strategic measures induced by admissible randomized policies is compact under the
weak convergence topology, if Assumption 5.2.1holdssothatT(dxt+1 |xt = x, ut = u) is weakly continuous in x, u and
also X, U are compact.

An implication of this result is that optimal policies exist, and are deterministic when the cost function is continuous in
x, u.
We note also that Schäl [283] introduces a more general topology, w − s topology, which requires strong continuity in
control actions. In this case, one can generalize Theorem 5.4.3 to the setups where Condition 2 applies and existence
of optimal policies follows.
We refer the reader to the Appendix, Section D.4, for a definition of the w-s topology.

Theorem 5.4.4 [283] The set of strategic measures induced by admissible randomized policies is sequentially compact
under the w-s topology, if T (dxt+1 |xt = x, ut = u) is strongly continuous in u for every x and also X, U are compact.

Thus, the above is a counterpart for when Assumption 5.2.2 holds.


The proofs of Theorems 5.4.3 and 5.4.4 follow from the property that to check whether a conditional independence
property, as in (5.18-5.19)), holds testing these on continuous and bounded functions implies this property for any
measurable and bounded function. Note that (5.19) holds since there is no conditional independence property condition,
and the main issue is to establish that (5.18) holds for any converging sequence of strategic measures. Applying the
hypotheses for each of the theorems leads to the desired results.
An implication of Theorem 5.4.4 is that an optimal strategic measure exists under the conditions of the theorem,
provided that the R+ -valued cost function c is lower semi-continuous in u for every x. In particular, for any w-s
converging sequence of strategic measures which satisfies (5.18)-(5.19), so does the limit. By [283, Theorem 3.7], and
the generalization of Portmanteau theorem for the w-s topology, the lower semi-continuity of the integral cost over the
set of strategic measures leads to the existence of an optimal strategic measure.
Now, we know that an optimal policy will be deterministic as a consequence of Theorem 5.4.1. Thus, an optimal policy
(which is deterministic) exists.

5.5 Infinite Horizon Optimal Discounted Cost Control Problems

When the time horizon becomes unbounded, we cannot directly invoke dynamic programming in the form considered
earlier. Infinite horizon problems that we will consider will belong to two classes: Discounted cost and average cost
problems. In the following, we first discuss the discounted cost problem. The average cost problem is discussed in
Chapter 7.
5.5 Infinite Horizon Optimal Discounted Cost Control Problems 95

Under the discounted cost criterion, future cost realizations are discounted: the future is perceived to be less important
than the current time with different justifications depending on the applications, e.g. due to the uncertainty in the
future leading one become more cautious about optimizing for the distant time stages, or perhaps due to an economic
understanding that the current value of a good is more important than its value in the future.
For a given T ∈ Z+ , the expected discounted cost criterion is given as:
TX−1
JβT (x0 , γ) = γ
Ex0 [ β t c(xt , ut )], (5.20)
t=0

for some β ∈ (0, 1). If there exists a policy γ ∗ which minimizes this cost, the policy is said to be optimal. We often
consider an infinite horizon problem by taking the limit (when c is non-negative)
T
X −1
Jβ (x0 , γ) = lim Exγ0 [ β t c(xt , ut )],
T →∞
t=0

and invoking the monotone convergence theorem:

X∞
Jβ (x0 , γ) = Exγ0 [ β t c(xt , ut )]. (5.21)
t=0

We seek to find
Jβ (x0 ) = inf Jβ (x0 , γ).
γ∈ΓA

Define
inf JβT (x0 , γ) = JβT (x0 )
γ∈ΓA

Lemma 5.5.1 Let A be a set and {fn } be a sequence of maps from fn : A → R for all n ∈ N. Then,

lim sup inf fn (x) ≤ inf lim sup fn (x).


n→∞ x∈A x∈A n→∞

Proof. For any n ∈ N and y ∈ A we have


inf fn (x) ≤ fn (y).
x∈A

This holds for all n we can take the limit superior of both sides, which yields

lim sup inf fn (x) ≤ lim sup fn (y).


n→∞ x∈A n→∞

This inequality holds for all y ∈ A and thus

lim sup inf fn (x) ≤ inf lim sup fn (x).


n→∞ x∈A x∈A n→∞


By Lemma 5.5.1, we change the order of limit and infimum so that

Jβ (x0 ) ≥ lim sup JβT (x0 ) (5.22)


T →∞

but since lim exists for the right-hand side as the expression is monotonically increasing the limit superior becomes an
actual limit and thus
Jβ (x0 ) ≥ lim JβT (x0 ).
T →∞

We will make use of this relation explicitly in Lemma 5.5.4 below. Now, observe that (from (5.20))
96 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming
 T
X −1 
JβT (x0 , γ) = Exγ0 c(x0 , u0 ) + E γ [ β t c(xt , ut )|x1 , x0 , u0 ] x0 , u0 ,
t=1

writes as
 TX−1 
JβT (x0 , γ) = Exγ0 c(x0 , u0 ) + βE γ [ β t−1 c(xt , ut )|x1 , x0 , u0 ] x0 , u0 .
t=1

Through the controlled Markov property and the fact that without any loss Markov policies are as good as any other
for finite horizon problems, it follows that (if dynamic programming recursions are well-defined)
 
T γ γ T −1
Jβ (x0 ) = inf Ex0 c(x0 , u0 ) + βE [Jβ (x1 )|x0 , u0 ] (5.23)
u0

We also saw in fact, under measurable selection conditions, via Bellman’s Theorem 5.1.3, the above is in fact an
equality. The goal is now to take T → ∞ and obtain desirable structural properties. The limit

lim JβT (x0 )


T →∞

will be a lower bound to Jβ (x0 ) by (5.22). But the inequality will turn out to be an equality under mild conditions to
be studied in the following. The next result is on the exchange of the order of the minimum and limits.

Lemma 5.5.2 [164] Let Vn (x, u) ↑ V (x, u) pointwise. Suppose that Vn and V are continuous in u for every x, and
u ∈ U(x) = U is compact. Then,
lim min Vn (x, u) = min V (x, u)
n→∞ u∈U(x) u∈U(x)

Proof. The proof follows from essentially the same arguments as in the proof of Lemma 5.2.2. Let u∗n solve
minu∈U(x) Vn (x, u). Note that

| min Vn (x, u) − min V (x, u)| ≤ V (x, u∗n ) − Vn (x, u∗n ), (5.24)
u∈U(x) u∈U(x)

since Vn (x, u) ↑ V (x, u). Now, suppose that for some ϵ > 0

V (x, u∗n ) − Vn (x, u∗n ) ≥ ϵ, (5.25)

along a subsequence nk . There exists a further subsequence n′k such that u∗n′ → ū for some ū. By assumption, for this
k
x and ū, and every ϵ > 0, we can find a sufficiently large N such that V (x, ū) − VN (x, ū) ≤ ϵ/2. Fix such an N . Now,
for every n′k ≥ N , since Vn is monotonically increasing:

V (x, u∗n′ ) − Vn′k (x, u∗n′ ) ≤ V (x, u∗n′ ) − VN (x, u∗n′ )


k k k k

However, V (x, u∗n′ ) and for the fixed N , VN (x, u∗n′ ), are continuous hence these two terms converge to: V (x, ū) −
k k
VN (x, ū). Hence (5.25) cannot hold. ⋄
Recall from dynamic programming equations that with

inf JβT (x0 , γ) = JβT (x0 ),


γ∈ΓA

we have (5.23):  
JβT (x0 ) = min c(x0 , u0 ) + βE γ
[JβT −1 (x1 )|x0 , u0 ] .
u0

It follows then that


 
Jβ∞ (x0 ) := lim JβT (x0 ) = lim min c(x0 , u0 ) + βE[JβT −1 (x1 )|x0 , u0 ] ,
T →∞ T →∞ u0
5.5 Infinite Horizon Optimal Discounted Cost Control Problems 97

where the limit exists due to the monotone convergence theorem since the cost is increasing with T : JβT (x1 ) ↑ Jβ∞ as
T → ∞. If Lemma 5.5.2 applies (i.e., the continuity condition in actions holds), we obtain that
 
∞ T
Jβ (x0 ) = min lim c(x0 , u0 ) + βE[Jβ (x1 )|x0 , u0 ] , (5.26)
u0 T →∞

and thus
 
Jβ∞ (x0 ) = min c(x0 , u0 ) + βE[Jβ∞ (x1 )|x0 , u0 ] . (5.27)
u0

The following result shows that the fixed point equation (5.27) is closely related to optimality. Define T as follows:
 Z 
(T(v))(x) := min c(x, u) + β v(y)T (dy|x, u) .
u X

and define the Discounted Cost Optimality Equation (DCOE) as follows

v(x) = (T(v))(x), x∈X (5.28)

Lemma 5.5.3 [Verification Theorem] [164]


(i) If v is a measurable R+ -valued function under Assumption 5.2.2 (or continuous and bounded function under
Assumption 5.2.1) with v ≥ Tv, then v(x) ≥ Jβ (x).
(ii) If Tv ≥ v and

lim β n Exγ [v(xn )] = 0, (5.29)


n→∞

for every policy and initial condition, then v(x) ≤ Jβ (x). As a result, a fixed point to (5.27) leads to an optimal
policy under (5.29).

Proof.
(i) For some stationary policy f that achieves (whose existence is justified by the measurable selection conditions)

min(c(x, u) + βE[v(x1 )|x0 = x, u0 = u]) = c(x, f (x)) + βE[v(x1 )|x0 = x, u0 = f (x)],

apply repeatedly
Z n−1
X
v(x) ≥ c(x, f (x)) + β v(y)T (dy|x, f (x)) ≥ · · · ≥ Exf [ β t c(xk , f (xk ))] + β n Exf [v(xn )]
k=0

Thus, taking the limit and given that v is non-negative valued,


n−1
X n−1
X
v(x) ≥ lim sup Exf [ β t c(xk , f (xk ))] + β n Exf [v(xn )] ≥ lim Exf [ β t c(xk , f (xk ))] ≥ Jβ (x),
n→∞ n→∞
k=0 k=0

since {f, f, f, · · · , f, · · · } is a particular policy and Jβ (x) is the optimal expected cost among all admissible poli-
cies.
(ii) If Tv(x) ≥ v(x), then for any n ∈ Z+ ,

Exγ [β n+1 v(xn+1 )|hn ] = Exγ [β n+1 v(xn+1 )|xn , un ]


 Z 
= β n {c(xn , un ) + β v(z)T (dz|xn , un )} − c(xn , un )
98 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

≥ β n (v(xn ) − c(xn , un )), (5.30)


R
where we use the inequality c(xn , un ) + β v(z)T (dz|xn , un ) ≥ v(xn ). Thus, using the iterated expectations and
arranging the terms
n−1
X n−1
X
Exγ [ β n c(xn , un )] ≥ E[ E[β n v(xk ) − β n+1 v(xk+1 )|hk ]]
k=0 k=0

leading to
n−1
X
Exγ [ β n c(xn , un )] ≥ v(x) − β n Exγ [v(xn )]
k=0

If the last term on the right hand size converges to zero, then the result is obtained so that for any fixed policy, v
provides a lower bound on the value function. Taking the infimum over all admissible policies, the desired result
v(x) ≤ Jβ (x) is obtained.

We have the following refinement, where we do not need to check (5.29) for every policy.

Lemma 5.5.4 If
v(x) = lim JβT (x)
T →∞

is so that v = T(v) where


T(v)(x) = c(x, f (x)) + βE[v(x1 )|x0 = x, u0 = f (x)]
is such that with γ = {f, f, · · · } ∈ ΓS ,

lim β n Exγ [v(xn )] = 0, (5.31)


n→∞

then γ is optimal.

Proof. Equation (5.22) implies that Jβ (x) ≥ v(x) since v is the pointwise limit of the discounted cost functions as the
horizon increases. Now, since the stationary policy f achieves
 
v(x) = min c(x, u) + βE[v(x1 )|x0 = x, u0 = u] = c(x, f (x)) + βE[v(x1 )|x0 = x, u0 = f (x)],
u∈X

applying this repeatedly to v(x0 ), v(x1 ) and then up to v(xn−1 ) leads to


Z n−1
X
v(x) = c(x, f (x)) + β v(x1 )T (dx1 |x, f (x)) = · · · = Exγ [ β k c(xk , f (xk ))] + β n Exγ [v(xn )].
k=0

Taking the limit, we have



X
v(x) = Exγ [ β k c(xk , f (xk ))]
k=0

implying that γ = {f, f, · · · } is optimal. ⋄


An implication of the proof of the result above is that for any stationary policy γ = {f, f, · · · , f, · · · }, we have the
following equation:

Jβ (x, γ) = c(x, f (x)) + βE[Jβ (x1 , γ)|x0 = x, u0 = f (x)] (5.32)

provided that
lim β n E γ [Jβ (xn , γ)] = 0.
n→∞

This will be useful later on when we study numerical methods.


5.5 Infinite Horizon Optimal Discounted Cost Control Problems 99

A sufficient condition for (5.31) is that the cost function c is bounded, though this is certainly not necessary.

5.5.1 Value Iteration Algorithm and Regularity of Value Functions

By dynamic programming, the Bellman optimality recursion for every finite horizon T ∈ N be written as
 Z 
JtT (x) = T(Jt+1
T
)(x) = min c(x, u) + β Jt+1 T
(y)T (dy|x, u) , t = T − 1, T − 2, · · · , 0, (5.33)
u X

with
JTT (x) = 0.
This sequence will lead to a solution for a T -stage discounted optimal cost problem. In particular, if we define v0 := JTT ,
and vn := JTT−n , we obtain the recursions for n = 1, 2, · · · ,

vn+1 = T(vn )(x),

which will form the basis of a very important algorithm, known as the value iteration algorithm, to be presented below.
The following then is a consequence of Lemma 5.5.4.

Theorem 5.5.1 [Value Iteration Algorithm: General Cost Setup] Suppose the cost function is non-negative. Consider
the successive iteration
Z
vn (x) = min{c(x, u) + β vn−1 (y)T (dy|x, u)}, ∀x, n ≥ 1 (5.34)
u X

with v0 (x) = 0 for all x ∈ X. Then, vn is a monotonically non-decreasing sequence. If this sequence converges
pointwise to a function v where
Z
v(x) = c(x, f (x)) + β v(y)T (dy|x, f (x))

is such that with γ = {f, f, · · · }, (5.31) holds, then γ is optimal and v is the value function.

A sufficient condition for the iterations in (5.33) to converge is the following. Suppose that measurable selection condi-
tions apply so that the iterations are well defined for every n ∈ Z+ . Let there exist a policy which leads to a finite cost
for every initial state and that by dynamic programming the recursions for every T given in (5.33) hold. This sequence
will lead to a solution for a T -stage discounted cost problem. Since JtT (x) ≤ JtT +1 (x), if there exists some Jt∞ such
that JtT (x) ↑ Jt∞ (x), we could invoke Lemma 5.5.2 to argue that
Z
Jt∞ (x) = T(Jt+1

)(x) = min{c(x, u) + β Jt+1 ∞
(y)T (dy|x, u)}.
u X

Such a limit exists, by the monotone convergence theorem since Jt∞ (x) < ∞ due to the assumption that there exists a
policy leading to a finite cost for every initial state. Hence, a limit satisfying (5.28) indeed exists. If
Z

{c(x, u) + β Jt+1 (y)T (dy|x, u)}
X

and Z
T
{c(x, u) + β Jt+1 (y)T (dy|x, u)}
X
are continuous in u for every x and every T and t, by (5.22), a lower bound to an optimal solution will have to satisfy
a fixed point equation (5.28). The result then would follow from Lemma 5.5.4.
In the bounded cost case, we can obtain a very strong result with a direct argument.
100 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

Lemma 5.5.5 (i) The space of measurable functions X → R endowed with the ||.||∞ norm (also called the supremum
norm) is a Banach space, that is

L∞ (X; R) = {f : X → R : ||f ||∞ = sup |f (x)| < ∞}


x

is a Banach space.
(ii) The space of continuous and bounded functions from X → R, Cb (X), endowed with the ||.||∞ norm is a Banach
space.

Theorem 5.5.2 [Value Iteration Algorithm - Bounded Cost Setup] Suppose the cost function is bounded, non-negative,
and one of the measurable selection conditions (Condition WF in Assumption 5.2.1 or Condition S in Assumption 5.2.2)
applies. Then, there exists a unique solution to the discounted cost problem which solves the fixed point equation.
Z
v(x) = min{c(x, u) + β v(y)T (dy|x, u)}, x∈X
u X

Furthermore, the optimal cost (value function) is obtained by a successive iteration (known as the Value Iteration
Algorithm):
Z
vn (x) = min{c(x, u) + β vn−1 (y)T (dy|x, u)}, ∀x, n ∈ N (5.35)
u X

For any v0 ∈ L∞ (X; R), the sequence converges to a unique fixed point. If v0 (x) = 0, x ∈ X, then vn (x) ↑ v(x) for
all x ∈ X (that is, vn monotonically converges to v). If Condition WF applies, then v is also continuous.

Proof of Theorem 5.5.2 Depending on the measurable selection conditions, we can take the value functions to be either
measurable and bounded, or continuous and bounded. (i) Suppose that we consider the measurable and bounded case
(Assumption 5.2.2). We observe that the vector J ∞ lives in L∞ (X; R) (since the cost is bounded, there is a uniform
bound for every x). We will show that the iteration given by
Z
T(v)(x) = min{c(x, u) + β v(y)T (dy|x, u)}
u X

is a contraction in L∞ (X; R). Let

||T(v) − T(v ′ )||∞ = sup |T(v)(x) − T(v ′ )(x)|


x∈X
Z Z
= sup min{c(x, u) + β v(y)T (dy|x, u)} − min{c(x, u) + β v ′ (y)T (dy|x, u)}
x∈X u X u X
  Z Z 
∗ ∗ ∗ ′ ∗
≤ sup 1{x∈A1} c(x, ux ) + β v(y)T (dy|x, ux ) − c(x, ux ) − β v (y)T (dy|x, ux )
x∈X X X
 Z Z 
∗∗ ∗∗ ∗∗ ′ ∗∗
+1{x∈A2} − c(x, ux ) − β v(y)T (dy|x, ux ) + c(x, ux ) + β v (y)T (dy|x, ux )
X X
 Z   Z 
′ ∗ ′ ∗∗
= sup 1{x∈A1} {β (v(y) − v(y ))T (dy|x, ux )} + sup 1{x∈A2} {β (v (y) − v(y))T (dy|x, ux )}
x∈X X x∈X X
Z Z
≤ β||v − v ′ ||∞ {1{x∈A1} T (dy|x, u∗∗x ) + 1{x∈A2} T (dy|x, u∗x )}
X X
= β||v − v ′ ||∞ (5.36)

Here  Z Z 
A1 = x : min{c(x, u) + β v(y)T (dy|x, u)} ≥ min{c(x, u) + β v ′ (y)T (dy|x, u)} ,
u X u X
∗∗ ∗
R
and A2 denotes the complementary
R ′ event, ux is a minimizing control for {c(x, u) + β X v(y)T (dy|x, u)} and ux is
a minimizer for {c(x, u) + β X v (y)T (dy|x, u)}. As a result T defines a contraction on the Banach space L∞ (X; R),
5.5 Infinite Horizon Optimal Discounted Cost Control Problems 101

and there exists a unique fixed point. Thus, the sequence of iterations in (5.33),
Z
T T T
Jt (x) = T(Jt+1 )(x) = {min{c(x, u) + β Jt+1 (y)T (dy|x, u)}},
u X

T
converges to J∞ (x) = J0∞ (x).
In particular, if one lets v0 (x) = 0 for all x ∈ X, the iterations increase monotonically and converges to the value
function. If one is only interested in convergence (and not the monotone behaviour), any initial function v0 ∈ L∞ (X; R)
is sufficient.
(ii) The above discussion also applies by considering a contraction on the space Cb (X), if Condition WF (Assumption
5.2.1) holds; in this case, the value function sequence vn is continuous for every n ∈ Z+ , and by the completeness of
Cb (X) under the supremum norm, so is the limit. ⋄

Example 5.4. Consider a controlled Markov chain with state space X = {0, 1}, action space U = {0, 1}, and transition
kernel for t ∈ Z+ :
P (xt+1 = 1|xt = 0, ut = 1) = P (xt+1 = 1|xt = 1, ut = 1) = α
P (xt+1 = 1|xt = 0, ut = 0) = P (xt+1 = 1|xt = 1, ut = 0) = 1 − α.
where α ∈ (0, 1). Let a cost function c(x, u), with c : X × U → R+ be given by

c(0, 1) = c(0, 0) = 1 c(1, 0) = c(1, 1) = 2.

Suppose that the goal is to minimize the quantity



X
E0γ [ β t c(xt , ut )],
t=0

for a fixed β ∈ (0, 1), over all admissible policies γ ∈ ΓA . Find an optimal policy and the optimal expected cost
explicitly, as a function of α, β (note that the initial condition is x0 = 0).

Solution. Apply value iteration. Take v0 ≡ 0. Then,

v1 (0) = min(c(0, u) + βE[v0 (x1 )|x0 = 0, u0 = u]) = 1


u

v1 (1) = min(c(1, u) + βE[v0 (x1 )|x0 = 1, u0 = u]) = 2


u

and

v2 (0) = min(c(0, u) + βE[v1 (x1 )|x0 = 0, u0 = u])


u
 
= min c(0, 0) + β(αv1 (0) + (1 − α)v1 (1)), c(0, 1) + β((1 − α)v1 (0) + αv1 (1))
u

v2 (1) = min(c(1, u) + βE[v1 (x1 )|x0 = 1, u0 = u])


u
 
= min c(1, 0) + β(αv1 (0) + (1 − α)v1 (1)), c(1, 0) + β((1 − α)v1 (0) + αv1 (1))
u

We see that if α < 12 , then the optimal selection is u = 1, leading to

v2 (0) = c(0, 1) + β((1 − α)v1 (0) + αv1 (1))


v2 (1) = c(1, 1) + β((1 − α)v1 (0) + αv1 (1))
102 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

We see that state 0 is always more desirable than state 1, in that v2 (0) < v2 (1) and this holds for all time stages. Then,
vn (0) ↑ v(0) and vn (1) ↑ v(1) with v(0) < v(1). As a result, for the discounted cost optimality equation

v(0) = min(c(0, u) + βE[v(x1 )|x0 = 0, u0 = u])


u
v(1) = min(c(1, u) + βE[v(x1 )|x1 = 0, u0 = u])
u

using v(0) < v(1), we have that:

v(0) = c(0, 1) + β((1 − α)v(0) + αv(1))

v(1) = c(1, 1) + β((1 − α)v(0) + αv(1))


Noting that v(1) = v(0) + 1 and solving for v, we obtain

1 + βα
v(0) = .
1−β

1
For α ≥ 2 a parallel argument can be made. ⋄

5.5.2 Lipschitz Regularity of Value Functions and the Case with Unbounded Costs

Lipschitz regularity of value functions

A further regularity property is the following. In the following W1 is the Wasserstein metric on probability measures;
see Appendix D. The following property will be useful later, when we study approximation and learning theoretic
applications.

Assumption 5.5.1 [Condition WA] Let d(·, ·) denote the metric on X. We assume that for some K1 , K2 :
(a) |c(x, u) − c(y, u)| ≤ K1 d(x, y); that is, c(·, u) is K1 -Lipschitz (denoted with the notation c(·, u) ∈ Lip(X, K1 )).
(b) W1 (T (dx1 |x0 = x, u0 = u), T (dx1 |x0 = y, u0 = u)) ≤ K2 d(x, y).

Theorem 5.5.3 [169] [270, Theorem 4.37] Suppose that Assumptions 5.2.1 and 5.5.1 hold, and βK2 < 1. Then, the
solution to
v = T(v)
K1
is Lipschitz with coefficient K = 1−βK2 .

Proof Let f ∈ Lip(X, k), that is f be k-Lipschitz for some k ∈ R. Then,



| Tf (z) − Tf (y) | ≤ max |c(z, u) − c(y, u)|
u∈U
Z Z 
+β f (x1 )η(dx1 |z, u) − f (x1 )η(dx1 |y, u)
X X
≤ K1 d(z, y) + βkK2 d(z, y) = (K1 + βkK2 ) d(z, y) =: M1 d(z, y). (5.37)

By induction we have for all n ≥ 2


Tn f ∈ Lip (X, Mn ) ,
Pn−1 i n K1
where Mn = K1 + βK2 Mn−1 and thus Mn = K1 i=0 (βK2 ) + k (βK2 ) . Taking k ≤ 1−βK 2
, we certify that the
K1
fixed point satisfies the desired Lipschitz continuity, as the sequence Mn monotonically converges to 1−βK 2
. Hence,
 
n K1
T f ∈ Lip X, 1−βK2 for all n, and therefore, the unique solution to the fixed point equation satisfies
5.5 Infinite Horizon Optimal Discounted Cost Control Problems 103
 
K1
v ∈ Lip X, (5.38)
1 − βK2

When X is not compact, note that we may still need to verify (5.31) to claim optimality.

Remark 5.5. We finally note that similar contraction arguments can also be applied to functions that are not necessar-
ily continuous, but only lower semi-continuous bounded functions, which also constitute a Banach space under the
supremum norm.

A further contraction argument for unbounded costs

As discussed earlier, one could follow the iteration method for the unbounded case (as in the proof of Theorem 5.5.1),
whereas the contraction method in the proof of Theorem 5.5.2 holds for the bounded cost case. The contraction method
can also be adjusted for the unbounded case under further conditions: If the cost is not bounded, one can define a
weighted sup-norm (called an f -norm): ∥c∥f = supx | fc(x)(x) |, where f is a positive function uniformly bounded from
below by a positive number. The contraction discussion above will apply to this context with such a consideration,
provided that the value function v used in the contraction analysis can be shown to satisfy ∥v∥f < ∞. For a suitable
function w, let Bw (X) denote the Banach space of measurable functions with a bounded w-norm. We state the corre-
sponding results formally in the following. We state two sets of conditions, one corresponds an unbounded function
generalization of strong continuity and the other of weak continuity conditions.

Assumption 5.5.2 (i) The one stage cost function c(x, u) is nonnegative and continuous in u for every x.
R
R stochastic kernel T ( · |x, u) is strongly continuous in u for every x, i.e., if uk → u, then u(y)T (dy|x, uk ) →
(ii) The
u(y)T (dy|x, u) for every measurable and bounded function u.
(iii)U is compact.
(iv) There exist nonnegative real numbers M and α ∈ [1, β1 ), and a weight function w : X → [1, ∞) such that for each
z ∈ X, we have

sup |c(x, u)| ≤ M w(x), (5.39)


u∈U
Z
sup w(y)T (dy|x, u) ≤ αw(x), (5.40)
u∈U X
R
and X
w(y)T (dy|x, u) is continuous in u for every x.

Assumption 5.5.3 (i) The one stage cost function c(x, u) is nonnegative and continuous in (x, u).
(ii) The stochastic kernel T ( · |x, u) is weakly continuous in (x, u) ∈ X × U, i.e., if (xk , uk ) → (x, u), then
T ( · |xk , uk ) → T ( · |x, u) weakly.
(iii)U is compact.
(iv) There exist nonnegative real numbers M and α ∈ [1, β1 ), and a continuous weight function w : X → [1, ∞) such
that for each z ∈ X, we have

sup |c(x, u)| ≤ M w(x), (5.41)


u∈U
Z
sup w(y)T (dy|x, u) ≤ αw(x), (5.42)
u∈U X
R
and X
w(y)T (dy|x, u) is continuous in (x, u).
104 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

Define the operator T on the set of real-valued measurable functions on X as


 Z 
Tv(z) = min c(z, a) + β v(y)T (dy|z, a) . (5.43)
a∈U X

It can be proved that T is a contraction operator mapping Bw (X) into itself with modulus σ = βα (see [165, Lemma
8.5.5]); that is,

∥Tu − Tv∥w ≤ β∥u − v∥w for all u, v ∈ Bw (X).

Theorem 5.5.4 [165, Theorem 8.3.6] [165, Lemma 8.5.5] Suppose Assumption 5.5.2 (or 5.5.3) holds. Then, the value
function J ∗ is the unique fixed point in Bw (X) (or Bw (X) ∩ C(X) under Assumption 5.5.3) of the contraction operator
T, i.e.,

J ∗ = TJ ∗ . (5.44)

Furthermore, a deterministic stationary policy f ∗ is optimal if and only if


Z
J ∗ (z) = c(z, f ∗ (z)) + β J ∗ (y)T (dy|z, f ∗ (z)). (5.45)
X

Finally, there exists a deterministic stationary policy f ∗ which is optimal (and thus satisfies (5.45)).

The proof follows from [165, Theorem 8.3.6]. See also [165, Lemma 8.5.5].

5.6 Regularity of Transition Kernels and Optimal Value Functions

We have seen in the chapter that continuity and regularity of transition kernels play a significant role for carrying out
optimality analysis. Later on we will see that these are also important for approximations, robustness, and learning
theoretic results and applications.
We review the following regularity properties for the transition kernels:
(i) T (·|x, u) is said to be weakly continuous (weak Feller) in (x, u), if T (·|xn , un ) → T (·|x, u) weakly for any
(xn , un ) → (x, u).
(ii) T (·|x, u) is said to be strongly continuous (strong Feller) in u for every x, if T (·|x, un ) → T (·|x, u) setwise for
any un → u for every fixed x ∈ X.
(iii)T (·|x, u) is said to be continuous under total variation in (x, u), if ∥T (·|xn , un ) − T (·|x, u)∥T V → 0 for any
(xn , un ) → (x, u).
(iv)T (·|x, u) is said to be continuous under the first order Wasserstein distance in (x, u), if

W1 (T (·|xn , un ), T (·|x, u)) → 0

for any (xn , un ) → (x, u). To ensure continuity of T with respect to the first order Wasserstein distance, in addition
to weak continuity, we may assume that there exists a function g : [0, ∞) → [0, ∞) such that as t → ∞, g(t) t ↑ ∞,
and Z
sup g(∥y∥) T (dy|x, u) < ∞
(x,u)∈K×U

for any compact K ⊂ X. Note that the latter condition implies uniform integrability of the collection of random vari-
ables with probability measures T (dx1 |X0 = xn , U0 = un ) as (xn , un ) → (x, u), which coupled with weak con-
vergence can be shown to imply convergence under the Wasserstein distance. To see this, let (xn , un ) → (x∞ , u∞ )
and Xn random variables with law T (·|xn , un ) satisfying Xn → X∞ a.s., which is possible by Skorohod’s theorem
5.6 Regularity of Transition Kernels and Optimal Value Functions 105

(Theorem B.3.5). Then the above condition implies uniform integrability of {Xn }, and thus E[∥Xn − X∞ ∥] → 0.
Then W1 (T (·|xn , un ), T (·|x∞ , u∞ )) → 0.

Example 5.6. Some example models satisfying these regularity properties are as follows:
(i) For a model with the dynamics xt+1 = f (xt , ut , wt ), the induced transition kernel T (·|x, u) is weakly continuous
in (x, u) if f (x, u, w) is a continuous function of (x, u), since for any continuous and bounded function g
Z Z
g(x1 )T (dx1 |xn , un ) = g(f (xn , un , w))µ(dw)
Z Z
→ g(f (x, u, w))µ(dw) = g(x1 )T (dx1 |x, u)

where µ denotes the probability measure of the noise process. If we also have that X is compact, the transition
kernel T (·|x, u) is also continuous under the first order Wasserstein distance.
(ii) For a model with the dynamics xt+1 = f (xt , ut ) + wt , the induced transition kernel T (·|x, u) is continuous under
total variation in (x, u) if f (x, u) is a continuous function of (x, u), and wt admits a continuous density function.
(iii)In general, if the transition kernel admits a continuous density function f so that T (dx1 |x, u) = f (x1 , x, u)dx1 ,
then T (dx1 |x, u) is continuous in total variation. This follows from an application of Scheffé’s Lemma [44, Theo-
rem 16.12]. In particular, we can write that
Z
∥T (·|xn , un ) − T (·|x, u)∥T V = |f (x1 , xn , un ) − f (x1 , x, u)|dx1 → 0.
X

(iv)For a model with the dynamics xt+1 = f (xt , ut , wt ), if f is Lipschitz continuous in (x, u) pair such that, there
exists some α < ∞ with

|f (xn , un , w) − f (x, u, w)| ≤ α (|xn − x| + |un − u|) ,

we can then bound the first order Wasserstein distance between the corresponding kernels with α:
Z Z
W1 (T (·|xn , un ), T (·|x, u)) = sup g(x1 )T (dx1 |xn , un ) − g(x1 )T (dx1 |x, u)
Lip(g)≤1
Z Z
= sup g(f (xn , un , w))µ(dw) − g(f (x, u, w))µ(dw)
Lip(g)≤1
Z
≤ |f (xn , un , w) − f (x, u, w)| µ(dw) ≤ α (|xn − x| + |un − u|) .

We next review the following regularity properties, which serves as a summary of Theorem 5.2.1, and Theorems 5.5.2
and 5.5.3 lead to the following.

Theorem 5.6.1 (Regularity for Finite Horizon Cost Criterion) Consider (5.1). Under Assumptions 5.2.1 and 5.2.2
there exists a minimizing control policy {ft , t ≥ 0} which is Markov (and thus in ΓM ). Furthermore, under Assumption
5.2.1, the function Jt , for all t, is continuous, under 5.2.2 Jt is Borel measurable, and under Assumptions 5.5.1 and
5.2.1 it is Lipschitz (in the latter case if cN exists, it is assumed to be Lipschitz).

Theorem 5.6.2 (Regularity for Discounted Cost Criterion) Consider (5.21). Under Assumptions 5.2.1 and 5.2.2
there exists a minimizing control policy which is stationary (and thus in ΓS ) without loss. Furthermore, under As-
sumption 5.2.1, the function Jβ is continuous, under 5.2.2, it is Borel measurable, and under Assumption 5.5.1 (with
βK2 < 1) and Assumption 5.2.1, it is Lipschitz
106 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

5.7 Exercises

Exercise 5.7.1 An investor’s wealth dynamics is given by the following:

xt+1 = ut wt ,

where {wt } is an i.i.d. R+ -valued stochastic process with E[ wt ] = 1 and ut is the investment of the investor at time
t. The investor has access to the past and current wealth information and his previous actions. The goal is to maximize:
T −1
X √
J(x0 , γ) = Exγ0 [ xt − ut ].
t=0

The investor’s action set for any given x is: U(x) = [0, x]. His initial wealth is given by x0 .
Formulate the problem as an optimal stochastic control problem by clearly identifying the state space, the control action
space, the information available at the controller at any time, the transition kernel and a cost functional mapping the
actions and states to R.
Find an optimal policy.
√ √
Hint: For α >√ 0, x − u√+ α u is a concave function of u for 0 ≤ u ≤ x and its maximum is computed when the
derivative of x − u + α u is set to zero.

Exercise 5.7.2 Consider the following linear system:

xt+1 = Axt + But + wt ,

where x ∈ Rn , u ∈ Rm and w ∈ Rn . Suppose {wt } is i.i.d. zero-mean Gaussian with a given covariance matrix
E[wt wtT ] = W for all t ≥ 0.
The goal is to obtain
inf J(x, γ),
γ

where
T
X −1
J(x, γ) = Exγ [ xTt Qxt + uTt Rut + xTT QT xT ],
t=0

with R, Q, QT > 0 (that is, these matrices are positive definite).


a) Show that there exists an optimal policy.
b) Obtain the Dynamic Programming recursion for the optimal control problem. Is the optimal control policy Markov?
Is it stationary?
c) For T → ∞, if (A, B) is controllable and with Q = C T C and (A, C) is observable, prove that the optimal policy
is stationary.

Exercise 5.7.3 (Optimality of Threshold Policies) [ [36]] Consider an inventory-production system given by

xt+1 = xt + ut − wt ,

where xt is R-valued, with the one-stage cost

c(xt , ut , wt ) = but + h max(0, xt + ut − wt ) + p max(0, wt − xt − ut )

Here, b is the unit production cost, h is the unit holding (storage) cost and p is the unit shortage cost; here we take
p > b. At any given time, the decision maker can take ut ∈ R+ . The demand variable wt ∼ µ is a R+ -valued i.i.d.
5.7 Exercises 107

process, independent of x0 , with a finite mean where µ is assumed to admit a probability density function. The goal is
to minimize
TX−1
J(x, γ) = Exγ [ c(xt , ut , wt )]
t=0

The controller at time t has access to It = {xs , us , s ≤ t − 1} ∪ {xt }.


Obtain a recursive form for the optimal solution. In particular, show that the solution is of threshold type: There exists
a sequence of real-numbers st so that the optimal solution is of the form: ut = 0 × 1{xt ≥st } + (st − xt ) × 1{xt <st } .
See [35] for a detailed analysis of this problem.
We note that the above is a relatively simplified model (e.g., the inventory can be negative and the cost/reward function
is simple) and such inventory problems can be made more general. Nonetheless, the solution method, via a convex
value function analysis and associated optimality arguments leading to threshold optimality, applies nearly identically
to many problems in stochastic control.

Exercise 5.7.4 (Optimal Stopping) [36]


Consider a burglar who is considering retirement. His goal is to maximize his earnings up to time T . At any time, he
can either continue his profession to steal an amount of wt which is an i.i.d. R+ -valued random process (he adds this
amount to his wealth), or retire.
However, each time he attempts burglary, there is a chance that he gets caught and he loses all of his savings (and
cannot work any further); this happens according to an i.i.d. Bernoulli process so that he gets caught with probability
p at each time stage when he is attempting to steal.
Assume that his initial wealth is x0 = 0. His goal is to maximize E[xT ]. Find his optimal policy for 0 ≤ t ≤ T − 1.
Note: Such problems where a decision maker can quit or stop a process are known as optimal stopping problems.

Exercise 5.7.5 (The Secretary Problem) Consider a manager who interviews N candidates for a position. The man-
ager wishes to maximize the probability of finding the best candidate. The candidates are interviewed in succession
according to a random order (uniformly distributed given all possible permutations). If a candidate is rejected, that
candidate is no longer available and if a candidate is selected, the process is over. The decisions must be made causally
given the available information up to that time, that is if the order is X1 , X2 , · · · , Xt , the policy can only use the in-
formation generated by σ(X1 , · · · , Xt ). What is the optimal strategy?
Hint: Apply dynamic programming. At time N , JN = N1 since at time N the past is given and the best one can
hope for is that the best candidate is the final one, which happens with probability N1 . Now, consider m = N − 1:
JN −1 = N 1−1 (max( NN−1 ), JN ) + NN −2
−1 JN . Here, the first event is the probability that the N − 1st candidate is the best
among the first N − 1 candidates, and in this case the manager needs to decide to stop or wait for the future (which
he does by comparing whether the best among the first N − 1 is the best among all, or whether he should skip and
N −2
move to the next time stage, N ). The second event is with probability N −1 in which case the best among the first N − 1
candidates is not the N − 1st one, in which case the manager has to wait. Continuing on with this logic:
1 m m−1
Jm = max( , Jm+1 ) + Jm+1
m N m
m
where N is the probability that the best in the first m is the best among all. Now, computing explicitly, we obtain

N −2 1 1
JN = 1/N ; JN −1 = ( + ); · · · ..
N N −2 N −1
1 1 1 m 1 1 1
and with Jm = m−1 m
N ( m−1 + m + · + N −1 ), and we stop when N ≥ Jm+1 = N ( m + m+1 + · + N −1 ). When
N is large enough, the above suggests that we stop as soon as the best candidate thus far has been spotted
PN −1 at time m
1 1
when 1 ≥ ( m + m+1 + · + N 1−1 ) ≈ log(N/m) which means that an optimal policy is around when m k1 < 1.
108 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

Approximately, this means that m∗ = N


e is a nearly optimal rule for large N if the m∗ th candidate is the best candidate
seen until then.

Exercise 5.7.6 A fishery manager annually has xt units of fish and sells ut xt of these where ut ∈ [0, 1]. With the
remaining ones, the next year’s production is given by the following model

xt+1 = wt xt (1 − ut ) + vt ,

with x0 is given and {wt , vt } is a sequence of mutually independent, identically distributed sequence of random vari-
ables with wt ≥ 0, vt ≥ 0 for all t and therefore E[wt ] = w̄ ≥ 0 and E[vt ] = v̄ > 0.
At time T , he sells all of the fish. The goal is to maximize the profit over the time horizon 0 ≤ t ≤ T − 1.
a) Formulate the problem as an optimal stochastic control problem by clearly identifying the state, the control actions,
the information available at the controller, the transition kernel and a cost functional mapping the actions and states
to R.
b) Does there exist an optimal policy? If it does, compute the optimal control policy as a dynamic programming
recursion.

Exercise 5.7.7 A common example in mathematical finance applications is the portfolio selection problem where a
controller (investor) would like to optimally allocate his wealth between a stochastic stock market and a market with a
guaranteed income : Consider a stock with an i.i.d. random return σt and a bank account with fixed interest rate r > 0.
These are modeled by:
Xt+1 = Xt ut (1 + σt ) + Xt (1 − ut )(1 + r), X0 = 1
and
Xt+1 = Xt (1 + r + ut (σt − r))
Here, ut ∈ [0, 1] denotes the proportion of the money that the investor invests in the stock market. Suppose that the
goal is to maximize E[log(XT )]. Then, we can write:
−1
TY T −1
Xk+1 X
log(XT ) = log( )= log((1 + r + ut (σt − r))) (5.46)
Xk
k=0 k=0

Formulate the problem as an optimal stochastic control problem by clearly identifying the state and the control action
spaces, the information available at the controller, the transition kernel, and a cost functional mapping the actions and
states to R. Find the optimal policy.

Exercise 5.7.8 We will illustrate dynamic programming by considering a simplified version of a setup in [155]. Con-
sider a two server-station network; where a router routes the incoming traffic, as is depicted in Figure 5.1.

Station 1

Station 2

Fig. 5.1
5.7 Exercises 109

Customers arrive according to a (continuous-time) Poisson process of rate λ. The router routes to station 1 with
probability u and second station with probability 1 − u. The router has access to the number of customers at both of
the queues, while implementing her policy.
Station 1 has a service time distribution which is exponential with rate µ1 , and Station 2 with µ2 = µ1 , as well. After
some computation, we find out that the controlled transition kernel is given by the following:

1 u
P (qt+1 = qt1 + 1, qt+1
2
= qt2 |qt1 , qt2 ) = λ
λ + 2µ1

1 (1 − u)
P (qt+1 = qt1 , qt+1
2
= qt2 + 1|qt1 , qt2 ) = λ
λ + 2µ1
1 µ1
P (qt+1 = max(qt1 − 1, 0), qt+1
2
= qt2 |qt1 , qt2 ) =
λ + 2µ1
1 µ1
P (qt+1 = qt1 , qt+1
2
= max(0, qt2 − 1)|qt1 , qt2 ) =
λ + 2µ1

There is also a holding cost per unit time. The holding cost at Station 1 is c1 > 0 and the cost at Station 2 is c2 > 0.
That is if there are qt1 customers, the cost is c1 qt1 at station 1 at time t and likewise for Station 2.
The goal of the router is to minimize the expected total holding cost from time 0 to some time T ∈ N, where the total
cost is
XT
c1 qt1 + c2 qt2 .
t=0

a) Express the problem as a dynamic programming problem, up until time T . That is; where does the control action
live? What is the state space? What is the transition kernel for the controlled Markov Chain?
Write down the dynamic programming recursion, starting from time T and going backwards.
b) Suppose that c1 = c2 . Let Jt (qt1 , qt2 ) be the value function at time t (that is the current cost and the cost to go).
Via dynamic programming, prove the following:
For a given t, if, whenever 0 ≤ qt1 ≤ qt2 we have that

Jt (qt1 , qt2 ) ≤ Jt (qt1 − 1, qt2 + 1),

then the same applies for Jt−1 (., .), for t ≥ 1. With the above, prove that an optimal control policy is given by:

ut = 1{qt1 ≤qt2 } ,

for all t values.

Exercise 5.7.9 Consider a scalar linear system with the following dynamics:

xt+1 = axt + but + wt ,

where {wt } is i.i.d Gaussian with zero-mean and unit variance. Suppose that the controller has access to It =
{x[0,t] , u[0,t−1] } at time t. Suppose that the initial state is x0 = x for some x ∈ R. We wish to find for some β ∈ (0, 1):

X
inf J(x0 , γ) = Exγ [ β t (qx2t + ru2t )],
γ
t=0

for q ≥ 0 and r > 0.


Compute the optimal control policy and the optimal cost.
110 5 Optimal Stochastic Control with Finite and Discounted Infinite Horizons and Dynamic Programming

Hint: Use Lemma 5.5.3. Start with a finite horizon version, and apply dynamic programming, obtain the solution and
take the finite horizon to infinity. This is also equivalent to applying value iteration with v0 (x) = 0 for all x ∈ R. You
will see that a recursion with vt (x) = Ct x2 + Dt will be obtained and Ct and Dt will have limits as t → ∞, C and
D, respectively. The optimal control will be stationary and deterministic:

ut = γ(xt ) = −(r + βCb2 )−1 βabCxt , t ≥ 0.

Thus, you need to find C and D.

Exercise 5.7.10 Consider a controlled Markov chain with state space X = {0, 1}, action space U = {0, 1}, and
transition kernel for t ∈ Z+ :

P (xt+1 = 1|xt = 0, ut = 1) = P (xt+1 = 1|xt = 0, ut = 0) = α

where α ∈ (0, 1). Furthermore,


1
P (xt+1 = 1|xt = 1, ut = 0) = P (xt+1 = 1|xt = 1, ut = 1) = .
2
Let a cost function c : X × U → R+ be given by

c(0, 1) = κ ∈ R+ , c(0, 0) = 1
1
c(1, 0) = , c(1, 1) = 1.
2

Suppose that the goal is to minimize the quantity



X
E0γ [ β t c(xt , ut )],
t=0

for a fixed β ∈ (0, 1), over all admissible policies γ ∈ ΓA .


Find an optimal policy and the optimal expected cost explicitly, as a function of α, β, κ.
6

Partially Observed Markov Decision Processes, Non-Linear Filtering, and


the Kalman Filter

As discussed earlier in Section 2.4, for a large class of problems the controller does not have access to the state process,
but may have access to some partial information obtained via noisy measurements. In particular, in this chapter, we
consider systems of the form:

xt+1 = f (xt , ut , wt ), yt = g(xt , vt ). (6.1)

Here, xt is the state (with x0 ∼ µ0 ), ut ∈ U is the control, (wt , vt ) are (W × V)-valued i.i.d noise processes where wt
is independent of vt .
The controller only has causal access to the second component {yt } of the process, together with the past applied
control actions. An admissible policy γ = {γt , t ∈ Z+ } is a collection of measurable functions so that γt is measurable
with respect to σ(It ) with It = {y[0,t] , u[0,t−1] } at time t. We emphasize the implicit assumption here that the control
policy can also (and typically does) depend on the prior probability measure µ0 . We denote the observed history space
as: H0 := Y, Ht = Ht−1 × Y × U.
In the following P(X) denotes the space of probability measures on X, which we assume to be a Polish space. Under
the topology of weak convergence, P(X) is also a Polish space (see Appendix D).

6.1 Enlargement of the State-Space and the Construction of a Controlled Markov Chain

We will see in this section that one could always transform a Partially Observable Markov Decision Problem (POMDP)
to a Fully Observed Markov Decision Problem (called a belief-MDP) via an enlargement of the state space and a re-
formulation of the model. In particular, when X, Y, U are countable (the more general case will be studied later in
the chapter in Section 6.3), we obtain via the properties of total probability the following recursion for conditional
probability measures, given an admissible policy,

P (xt = x, yt , ut−1 |y[0,t−1] , u[0,t−2] )


πt (x) := P (xt = x|y[0,t] , u[0,t−1] ) = P
x∈X P (xt = x, yt , ut−1 |y[0,t−1] , u[0,t−1] )
γ
P
x ∈X P (yt |xt )P (xt |xt−1 , ut−1 )P (ut−1 |y[0,t−1] , u[0,t−2] )πt−1 (xt−1 )
= P Pt−1 γ
.
xt−1 ∈X x∈X P (yt |xt = x)P (xt = x|xt−1 , ut−1 )P (ut−1 |y[0,t−1] , u[0,t−2] )πt−1 (xt−1 )
P
x ∈X P (yt |xt )P (xt |xt−1 , ut−1 )πt−1 (xt−1 )
= P Pt−1 .
xt−1 ∈X x∈X P (yt |xt = x)P (xt = x|xt−1 , ut−1 )πt−1 (xt−1 )
=: F (πt−1 , yt , ut−1 )(x) (6.2)

Notice that the right hand side does not depend on the policy γ, therefore the conditional expectation is policy-
independent. We will see shortly that the conditional measure process forms a controlled Markov chain in P(X).
112 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

Note that in the above analysis P (ut−1 |y[0,t−1] , u[0,t−2] ) is determined by the control policy, and P (xt |xt−1 , ut−1 ) is
determined by the transition kernel T of the controlled Markov chain.
The result above leads to the following.

Theorem 6.1.1 The process {πt , ut } is a controlled Markov chain. That is, under any admissible control policy, given
the action at time t ≥ 0 and πt , πt+1 is conditionally independent from {πs , us , s ≤ t − 1}.

We will prove the result for the case where Y is countable. For the more general case, see Section 6.3.
Proof. Let D ∈ B(P(X)). From (6.2), with Yy+1 denoted with the capital letter to emphasize its randomness, we have
under any admissible policy,

P (πt+1 ∈ D|πs , us , s ≤ t) = P (F (πt , Yt+1 , ut ) ∈ D|πs , us , s ≤ t)


= P (F (πt , Yt+1 , ut ) ∈ D, Yt ∈ Y|πs , us , s ≤ t)
X
= P (F (πt , yt+1 , ut ) ∈ D, yt+1 = y|πs , us , s ≤ t)
y∈Y
X
= P (F (πt , yt+1 , ut ) ∈ D|yt+1 = y, πs , us , s ≤ t)P (yt+1 = y|πs , us , s ≤ t)
y∈Y
X
= 1  P (yt+1 = y|πt , ut )
y∈Y F (πt ,y,ut )∈D

X XX 
′ ′
= 1  P (yt+1 = y|xt+1 = x )P (xt+1 = x |xt = x, ut )πt (x)
y∈Y F (πt ,y,ut )∈D x′ ∈X x∈X

= P (πt+1 ∈ D|πt , ut ) =: η(D|πt , ut ) (6.3)

Observe that the kernel η does not depend on the policy (and thus it is policy-independent). We still need to show that
the expression P (πt+1 ∈ ·|πt , ut ) : P(X) × U → P(P(X)) is a regular conditional probability measure; that is, for
every fixed D ∈ B(P(X)), P (πt+1 ∈ D|πt , ut ) is a measurable function on P(X) × U and for every πt , ut , the map is
a probability measure on P(X). The rest of the proof follows in Section 6.3. ⋄
Let the cost function to be minimized be
T
X −1
Eµγ0 [ c(xt , ut )],
t=0

where Eµγ0 [·] denotes the expectation over all sample paths with initial state measure given by µ0 under policy γ =
{γ0 , γ1 , · · · }. We can transform the system into a fully observed Markov model as follows. Using the law of the iterated
expectations (Theorem 4.1.3), write the total cost as

T
X −1 T
X −1
Eµγ0 [ c(xt , ut )] = Eµγ0 [ E γ [c(xt , ut )|It ]].
t=0 t=0

Given a policy γ with ut = γt (It ), we have that


T
X −1 −1
 TX  
Eµγ0 [ c(xt , ut )] = Eµγ0 γ
E c(xt , γt (It ))|It
t=0 t=0
−1  X X
 TX 
= Eµγ0 P γ (xt = x|It , ut )P γ (ut = u|It )c(x, u)
t=0 x∈X u∈U
−1  X
 TX 
= Eµγ0 P γ (xt = x|It )c(x, γt (It ))
t=0 x∈X
6.1 Enlargement of the State-Space and the Construction of a Controlled Markov Chain 113
−1  X
 TX 
= Eµγ0 πt (x)c(x, γt (It )) (6.4)
t=0 x∈X

Notice that P γ (xt = x|It , ut ) = P γ (xt = x|It ) = P (xt = x|It ) is policy-independent as noted earlier. At this
point, we should pause and reflect on Theorem 6.1.1 and (Blackwell’s) Theorem 5.1.1, to conclude that without any
loss a policy, for finite horizons, could use πt and t, by the following reasoning: Define a stage-wise cost function
c̃ : P(X) × U → R+ as
X
c̃(π, u) = c(x, u)π(x), π ∈ P(X), (6.5)
X

Observe that an admissible control policy will select ut as a function of It . However, we know that for a finite horizon
problem, by Blackwell’s Theorem 5.1.1, any admissible policy can be replaced with one which only uses πt without
any loss (since πt , ut forms a controlled Markov chain), and therefore without any loss, we can restrict our search space
to policies which are Markov (that is which only use πt and t).
In view of the preceding discussion, it follows then that an optimal solution to the following minimization for the
problem
TX−1
Eµγ0 [ c̃(πt , ut )],
t=0

for the controlled belief-MDP model with kernel η (6.3), is also optimal for the original problem (6.4), where the initial
state distribution µ0 for the belief-MDP is the probability measure on π0 (·) = P (x0 ∈ ·|y0 ) induced by the initial
probability measure µ0 on x0 and the measurement variable y0 .
Let η be the transition kernel defined with (6.3). It follows then that (P, U, η, c̃) defines a completely observable
controlled Markov process (also called a belief-MDP).
Thus, one can obtain the optimal solution by using the solution of the filtering equation (6.2) as a sufficient statistic,
as Markov policies (policies that use the Markov state as their sufficient statistics) are optimal for control of Markov
chains, under the previously studied measurable selection conditions (see Section 5.2) which require some regularity
conditions. We will discuss these later in the chapter.
We call the control policies which use π as their information to generate control as separated control policies; as one
first generates the belief πt via the filtering equation, and then generates the control via πt .
We note here that some of the first separation results for partially observed Markov Decision Processes were reported
in [352], [296], and [262], among others.
Separation will be particularly consequential in the context of linear Gaussian systems: A Gaussian probability measure
can be uniquely identified by knowing the mean and the covariance of the Gaussian random variable. This makes the
analysis for estimating a Gaussian random variable particularly simple to perform, since the conditional estimate of
a partially observed (through an additive Gaussian noise) Gaussian random variable is a linear/affine function of the
observed variable and the non-linear filtering equation (6.2) becomes significantly simpler. Recall that a Gaussian
measure with mean µ and covariance matrix KXX has the following density:
1 T −1
p(x) = n e−1/2((x−µ) KXX (x−µ)) ,
(2π) 2 |KXX |1/2

and thus it suffices to compute the mean and the covariance matrix to define the Gaussian probability measure.
114 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

6.2 The Linear Quadratic Gaussian (LQG) Problem and Kalman Filtering

6.2.1 A Supporting Result on Estimation

Lemma 6.2.1 Let X be a random variable (defined on a probability space (Ω, F, P )) with a finite second moment and
R > 0 (that is, a positive definite matrix). The following holds

inf E[(X − g(Y ))T R(X − g(Y ))] = E[(X − G(Y ))T R(X − G(Y ))],
g∈M(Y)

where M(Y) denotes the set of measurable functions from Y to R and where G(y) = E[X|Y = y] almost surely.

Before we state the proof, it is useful to emphasize that there are setups where the measurability assumption is not
superfluous. See Exercise 4.5.13.
Proof. Let G(y) = E[X|Y = y] + h(y), for some measurable h; we then have the following through the law of the
iterated expectations:

E[(X − E[X|Y ] − h(Y ))T R(X − E[X|Y ] − h(Y ))]


= E[(X − E[X|Y ])T R(X − E[X|Y ])] + 2E[(X − E[X|Y ])T Rh(Y )] + E[hT (Y )Rh(Y )]
= E[(X − E[X|Y ])T R(X − E[X|Y ])] + 2E[E[(X − E[X|Y ])T Rh(Y )|Y ]] + E[hT (Y )Rh(Y )] (6.6)
 
= E[(X − E[X|Y ])T R(X − E[X|Y ])] + E[hT (Y )Rh(Y )] + 2E E[(X − E[X|Y ])T |Y ]Rh(Y ) (6.7)

= E[(X − E[X|Y ])T R(X − E[X|Y ])] + E[hT (Y )Rh(Y )]


≥ E[(X − E[X|Y ])T R(X − E[X|Y ])],

where in (6.6) we use Theorem 4.1.3 and in (6.7) we use Theorem 4.1.4. Note that without any loss we can assume that
E[hT (Y )h(Y )] < ∞ (by the above analysis for otherwise the expectation above would be unbounded) and therefore
  1/2  1/2
T T T
E[|(X − E[X|Y ]) Rh(Y )|] ≤ E (X − E[X|Y ]) R(X − E[X|Y ]) E[h (Y )Rh(Y )]

by the Cauchy-Schwarz inequality, so that (X − E[X|Y ])T Rh(Y ) is integrable, validating the use Theorem 4.1.3.
Thus, for an optimal policy, we must have that E[hT (Y )Rh(Y )] = 0. ⋄

Remark 6.1. We note that the above admits a Hilbert space interpretation or formulation: Let H denote the space
of random variables (defined on a probability space) on which the inner product ⟨X, Z⟩ := E[X T RZ] is defined;
this defines a Hilbert space. Let MH be a subspace of H, defined as the closed subspace of random variables that
are measurable on σ(Y ). Then, the projection theorem [219] leads to the observation that an optimal g(Y ) ∈ MH
minimizing ∥X − g(Y )∥22 , denoted here by G(Y ), is one which satisfies:

⟨X − G(Y ), h(Y )⟩ = E[X − G(Y )T Rh(Y )] = 0, ∀h ∈ MH

The conditional expectation satisfies this condition as

E[(X − E[X|Y ])T Rh(Y )] = E[E[(X − E[X|Y ])T Rh(Y )|Y ]] = E[E[(X − E[X|Y ])T |Y ]Rh(Y )] = 0,

since P a.s., E[(X − E[X|Y ])T |Y ] = 0.

6.2.2 The Linear Quadratic Gaussian Problem

Consider the following linear system:


6.2 The Linear Quadratic Gaussian (LQG) Problem and Kalman Filtering 115

xt+1 = Axt + But + wt ,


yt = Cxt + vt , (6.8)

where x ∈ Rn , u ∈ Rm and w ∈ Rn , y ∈ Rp , v ∈ Rp . Suppose {wt , vt } are zero-mean i.i.d. random Gaussian vectors
with given covariance matrices E[wt wtT ] = W and E[vt vtT ] = V for all t ≥ 0.
The goal is to obtain
inf J(γ, µ0 ),
γ

where
N
X −1
J(µ0 , γ) = Eµγ0 [ xTt Qxt + uTt Rut + xTN QN xN ], (6.9)
t=0

with R > 0 and Q, QN ≥ 0 (that is, these matrices are positive definite and positive semi-definite) and µ0 is an initial
prior probability measure (on x0 ) assumed to be zero-mean Gaussian.
Building on Lemma 6.2.1, we will show in the following that the optimal control is linear in its expectation and has the
form
ut = −(B T Kt+1 B + R)−1 B T Kt+1 AE[xt |It ]
where Kt solves the Discrete-Time Riccati Equation:

Kt = Q + AT Kt+1 A − AT Kt+1 B(B T Kt+1 B + R)−1 B T Kt+1 A,

with final condition KN = QN .


In the following, we start with the estimation problem.

6.2.3 Estimation and Kalman Filtering

In this section, we discuss the control-free setup and derive the celebrated Kalman Filter. In the following to make
certain computations more explicit and easier to follow, we will use capital letters to denote the random variables and
small letters for the realizations of these variables.
For a linear Gaussian system, the state process has a Gaussian probability measure. A Gaussian probability measure
can be uniquely identified by knowing the mean and the covariance of the Gaussian random variable. This makes the
analysis for estimating a Gaussian random variable particularly simple to perform, since the conditional estimate of
a partially observed (through an additive Gaussian noise) Gaussian random variable is a linear/affine function of the
observed variable.
Recall that a Gaussian measure with mean µ and covariance matrix ΣXX has the following density:
1 T −1
p(x) = n e−1/2((x−µ) ΣXX (x−µ))
(2π) 2 |ΣXX |1/2

Lemma 6.2.2 Let X, Y be zero-mean Gaussian vectors. Then E[X|Y = y] is linear in y: With ΣXY = E[XY T ] and
ΣY Y = E[Y Y T ],

E[X|Y = y] = ΣXY ΣY−1Y y, (6.10)

and
E[(X − E[X|Y ])(X − E[X|Y ])T ] = ΣXX − ΣXY ΣY−1Y ΣXY
T
=: D.
In particular,
E[(X − E[X|Y ])(X − E[X|Y ])T |Y = y]
116 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

does not depend on the realization y of Y and is equal to D.

We note that if the random variables are not-zero mean, one needs to add a constant correction term making the estimate
an affine function (of the measurement).
  
X
Proof. By Bayes’ rule and the fact that the processes admit densities: p(x|y) = p(x,y) .With K XY := E [X T T
Y ] ,
p(y) Y
we have that  
ΣXX ΣXY
KXY =
ΣY X ΣY Y
−1
It follows that KXY is also symmetric (since the eigenvectors are the same as those of KXY and the eigenvalues are
inverted) and given by:  
−1 Ψ Ψ
KXY = XX XY ,
ΨY X ΨY Y
Thus, for some normalization constant C,
T T T
p(x, y) e−1/2(x ΨXX x+2x ΨXY y+y ΨY Y y)
=C −1
p(y) T
e−1/2(y KY Y y)
By the completion of the squares method for the expression in the exponent, for some matrix D we obtain

(xT ΨXX x + 2xT ΨXY y + y T ΨY Y y − y T KY−1Y y) = (x − Hy)T D−1 (x − Hy) + Q(y),


−1 −1 −1
it follows that H = −ΨXX ΨXY and D = ΨXX . Since KXY KXY = I (and thus ΨXX ΣXY + ΨXY ΣY Y = 0), H is
−1
also equal to ΣXY ΣY Y . Here Q(y) is a quadratic expression in y. As a result, one obtains
1 T
D −1 (x−Hy)
p(x|y) = Ce− 2 Q(y) e−1/2(x−Hy) .
1 1
p(x|y)dx = 1 (as it is a conditional probability density function), it follows that Ce− 2 Q(y) =
R
Since n and
(2π) 2 |D|1/2
is in fact independent of y. Then, we finally have that D, which does not depend on y, equals (see Lemma 6.2.4 below)

E[(X − E[X|Y ])(X − E[X|Y ])T ] = E[XX T ] − E[(E[X|Y ])(E[X|Y ])T ] = ΣXX − ΣXY ΣY−1Y ΣXY
T
(6.11)

Remark 6.2. The fact that Q(y) above does not depend on y reveals an interesting result that the conditional covariance
of X − E[X|Y ] viewed as a Gaussian random variable is identical for all y values. This is a crucial fact that will be
utilized in the derivation of the Kalman Filter.

Remark 6.3. Even if the random variables X, Y are not Gaussian (but zero-mean), through another Hilbert space formu-
lation and an application of the Projection Theorem (see Remark 6.1), it can be shown that the expression ΣXY ΣY−1Y y
is the best linear estimate, that is the solution to inf K E[(X − KY )T (X − KY )]. One can naturally generalize this for
random variables with non-zero mean.

We will derive the Kalman filter in the following. The following two lemmas are instrumental.

Lemma 6.2.3 If E[X] = 0 and Z1 , Z2 are orthogonal zero-mean Gaussian processes (with E[Z1 Z2T ] = 0), then
E[X|Z1 = z1 , Z2 = z2 ] = E[X|Z = z1 ] + E[X|Z2 = z2 ].

−1
Proof. The proof follows by writing z = [z1 , z2 ]T , noting that ΣZZ is diagonal and E[X|z] = ΣXZ ΣZZ z. ⋄

Lemma 6.2.4 E[(X − E[X|Y ])(X − E[X|Y ])T ] is given by D = ΣXX − ΣXY ΣY−1Y ΣXY
T
above.
6.2 The Linear Quadratic Gaussian (LQG) Problem and Kalman Filtering 117

Proof. Note that

E[X(E[X|Y ])T ] = E[(X − E[X|Y ] + E[X|Y ])(E[X|Y ])T ] = E[E[X|Y ](E[X|Y ])T ]

since X − E[X|Y ] is orthogonal to E[X|Y ]1 . As a result,

E[(X − E[X|Y ])(X − E[X|Y ])T ] = E[XX T ] − 2E[X(E[X|Y ])T ] + E[E[X|Y ](E[X|Y ])T ]
= E[XX T ] − E[E[X|Y ](E[X|Y ])T ],

and the result follows from Lemma 6.2.2 in view of (6.10). ⋄


Now, we can move on to the derivation of the Kalman Filter.
Consider
xt+1 = Axt + wt , yt = Cxt + vt ,
with E[wt wtT ] = W and E[vt vtT ] = V where {wt } and {vt } are mutually independent i.i.d. zero-mean Gaussian
processes.
Define
mt = E[xt |y[0,t−1] ]
Σt|t−1 = E[(xt − E[xt |y[0,t−1] ])(xt − E[xt |y[0,t−1] ])T |y[0,t−1] ]
and note that since the estimation error covariance does not depend on the realization y[0,t−1] (see Remark 6.2), we
write also
Σt|t−1 = E[(xt − E[xt |y[0,t−1] ])(xt − E[xt |y[0,t−1] ])T ]

Theorem 6.2.1 The following holds:

mt+1 = Amt + AΣt|t−1 C T (CΣt|t−1 C T + V )−1 (yt − Cmt ) (6.12)


T T T −1 T
Σt+1|t = AΣt|t−1 A + W − (AΣt|t−1 C )(CΣt|t−1 C + V ) (CΣt|t−1 A ) (6.13)

with
m0 = E[x0 ]
and
Σ0|−1 = E[x0 xT0 ]

Proof. With xt+1 = Axt + wt , the following hold:

mt+1 = E[Axt + wt |y[0,t] ] = E[Axt |y[0,t] ] = E[Amt + A(xt − mt )|y[0,t] ]


= Amt + E[A(xt − mt )|y[0,t−1] , yt − E[yt |y[0,t−1] ]]
= Amt + E[A(xt − mt )|y[0,t−1] ] + E[A(xt − mt )|yt − E[yt |y[0,t−1] ]] (6.14)
= Amt + E[A(xt − mt )|yt − E[yt |y[0,t−1] ]]
= Amt + E[A(xt − mt )|Cxt + vt − E[Cxt + vt |y[0,t−1] ]]
= Amt + E[A(xt − mt )|C(xt − mt ) + vt ] (6.15)

In the above, (6.14) follows from Lemma 6.2.3. We also use the fact that wt is independent of (and hence orthogonal
to) y[0,t] . Let X = A(xt − mt ) and Y = yt − E[yt |y[0,t−1] ] = yt − Cmt = C(xt − mt ) + vt . Then, by Lemma 6.2.2,
E[X|Y ] = ΣXY ΣY−1Y Y and thus,

  
Note that by iterated expectations, we have E[(X −E[X|Y ])(E[X|Y ])T ] = E E[(X −E[X|Y ])(E[X|Y ])T |Y ] = E E[(X −
1


T
E[X|Y ])|Y ](E[X|Y ]) = 0.
118 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

mt+1 = Amt + AΣt|t−1 C T (CΣt|t−1 C T + V )−1 (yt − Cmt )

Likewise,
xt+1 − mt+1 = A(xt − mt ) + wt − AΣt|t−1 C T (CΣt|t−1 C T + V )−1 (yt − Cmt ),
leads to, after a few lines of calculations:

Σt+1|t = AΣt|t−1 AT + W − (AΣt|t−1 C T )(CΣt|t−1 C T + V )−1 (CΣt|t−1 AT )


The above is the celebrated Kalman filter.
Define now
m̃t := E[xt |y[0,t] ] = mt + E[xt − mt |y[0,t] ]
Following the analysis above, we obtain
 
m̃t = mt + E[xt − mt |y[0,t−1] ] + E xt − mt |yt − E[yt |y[0,t−1] ] .

Note that we also have mt = Am̃t−1 . The following result then follows:

Theorem 6.2.2 The recursions for m̃t satisfy

m̃t = Am̃t−1 + Σt|t−1 C T (CΣt|t−1 C T + V )−1 (yt − CAm̃t−1 ), (6.16)

with m̃0 = E[x0 |y0 ].

We observe that the zero-mean variable xt − m̃t is orthogonal to y[0,t] , in the sense that the error is independent of the
information available at the controller, and since the information available is Gaussian, independence and orthogonality
are equivalent.
We observe that the recursion (6.13) in Theorem 6.2.1 is essentially identical to the recursions in Theorem 5.3.2 with
writing A = AT , W = Q, V = R, C T = B. This leads to the following result (as a corollary of Theorem 6.2.1).

Theorem 6.2.3 Suppose (AT , C T ) is controllable (this is equivalent to saying that (A, C) is observable) and V > 0.
Then, the recursions for the covariance matrices Σt in Theorem 6.2.1 admit a fixed point. If, in addition, with W =
BB T , (AT , B T ) is observable (that is (A, B) is controllable), the fixed point solution is unique, and is positive definite.
As noted earlier, these can be relaxed to stabilizability (of (A, B)) and detectability of (A, C) but in this case the fixed
point solution may only be positive semi-definite.

Remark 6.4. The above suggest that if the observations are sufficiently informative, then the Kalman filter converges to
a solution (with an appropriate initialization), even in the absence of an irreducibility condition (i.e., the controllability
condition for (A, B) above) on the original state process xt ; under irreducibility, however, the solution is unique. This
intuition has been shown to find a precise generalization in the non-linear filtering context [84,226,314], see Definition
6.4.10.

6.2.4 Optimal Control of Partially Observed LQG Systems

Let us revisit (6.9). With the analysis of optimal linear estimation above, we will now reformulate the quadratic opti-
mization problem (6.9) in terms of m̃t , ut and xt − m̃t as follows. First, let us note the following:

Theorem 6.2.4 Consider the controlled linear system (6.8). Then, with

mt = E[xt |y[0,t−1] , u[0,t−1] ]


6.2 The Linear Quadratic Gaussian (LQG) Problem and Kalman Filtering 119

and
Σt|t−1 = E[(xt − E[xt |y[0,t−1] , u[0,t−1] )(xt − E[xt |y[0,t−1] , u[0,t−1] ])T ],
the following hold:

mt+1 = Amt + But + AΣt|t−1 C T (CΣt|t−1 C T + V )−1 (yt − Cmt )


Σt+1|t = AΣt|t−1 AT + W − (AΣt|t−1 C T )(CΣt|t−1 C T + V )−1 (CΣt|t−1 AT )

with
m0 = E[x0 ]
and
Σ0|−1 = E[x0 xT0 ]

The proof follows that of Theorem 6.2.1: the only difference is the presence of control. Observe that, the estimation
can be viewed to be that of estimating:
 n−1
X  n−1
X n−1
X
n n−k−1
xn = A x0 + A wk + An−k−1 Buk =: x̄n + An−k−1 Buk
k=0 k=0 k=0

where
x̄n+1 = Ax̄n + wn
Pn−1 n−k−1
is the control-free system. But since k=0 A Buk is known at time n (by the controller), the estimation problem
is essentially that of estimating the control-free system x̄n . Furthermore, the control adds no additional information
with regard to estimating x̄n , that is, the information generated by

ȳn = C x̄n + vn

up to time n contains the same information with regard to x̄n as that contained by {yk , uk } up to time n, because
(i) x̄n is not affected by the control, and (ii) the information that control actions contain are already available in the
information content of the current and past ȳn variables under any measurable policy (to see this, note that u0 is a
function of ȳ0 , and u1 is a function of u0 and ȳ[0,1] , and thus really only that of ȳ[0,1] , and so on for n > 1 by an
inductive reasoning). That is, under any policy γ, for any Borel B and any n:

P γ (x̄n ∈ B|ȳ[0,n] ) = P γ (x̄n ∈ B|ȳ[0,n] , u[0,n−1] )

What the above implies is that, under any policy γ


n−1
X n−1
X
γ n−k−1
E [xn |y[0,n] , u[0,n−1] ] = A Buk + E[x̄n |y[0,n] , u[0,n−1] ] = An−k−1 Buk + E[x̄n |ȳ[0,n] ].
k=0 k=0

Furthermore, xn − E[xn |y[0,n−1] , u[0,n−1] ] is sample path equivalent to x̄n − E[x̄n |ȳ[0,n−1] ], and these are determined
solely by x0 , w[0,n−1] , v[0,n−1] .
Now, for the controlled case, let us define

m̃t = E[xt |y[0,t] , u[0,t−1] ].

and observe that the Kalman filtering recursions apply almost verbatim with the control actions added in an additive
fashion:
 
m̃t = Am̃t−1 + But−1 + Σt|t−1 C T (CΣt|t−1 C T + V )−1 yt − C(Am̃t−1 + But−1 ) (6.17)

Let It = {y[0,t] , u[0,t−1] }. Observe now that


120 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

E[xTt Qxt ] = E[(xt − m̃t + m̃t )T Q(xt − m̃t + m̃t )]


= E[(xt − m̃t )T Q(xt − m̃t )] + E[m̃Tt Qm̃t ] + 2E[(xt − m̃t )T Qm̃t ]
= E[(xt − m̃t )T Q(xt − m̃t )] + E[m̃Tt Qm̃t ] + 2E[E[(xt − m̃t )T Qm̃t |It ]]
= E[(xt − m̃t )T Q(xt − m̃t )] + E[m̃Tt Qm̃t ] (6.18)

by the orthogonality property of the conditional estimation error to m̃t (recall that m̃t is a function of It and E[(xt −
m̃t )T Qm̃t |It ] = 0). In particular, the cost:
NX−1
γ
J(γ, µ0 ) = Eµ0 [ xTt Qxt + uTt Rut + xTN QN xN ], (6.19)
t=0

writes as:
N
X −1 N
X −1
Eµγ0 [ m̃Tt Qm̃t + uTt Rut + m̃TN QN m̃N ] + Eµγ0 [ (xt − m̃t )T Q(xt − m̃t )]
t=0 t=0
γ
+Eµ0 [(xN − m̃N )T QN (xN − m̃N )] (6.20)

for the fully observed system (see (6.17)):

m̃t = Am̃t−1 + But−1 + w̃t−1 , (6.21)

with  
w̃t−1 = Σt|t−1 C T (CΣt|t−1 C T + V )−1 yt − C(Am̃t−1 + But−1 )

Furthermore, the estimation errors in (6.20) (the second and the third terms) do not depend on the control policy γ so
that the expected cost writes as
N
X −1 N
X −1
Eµγ0 [ m̃Tt Qm̃t + uTt Rut + m̃TN QN m̃N ] + Eµ0 [ (xt − m̃t )T Q(xt − m̃t )]
t=0 t=0
+Eµ0 [(xN − m̃N )T QN (xN − m̃N )] (6.22)

Thus, the optimal control problem is equivalent to the control of the fully observed state m̃t , with additive time-varying
independent Gaussian noise process {w̃t } given in (6.21).
Here, that the error term (xt − m̃t ) does not depend on the control policy is a consequence of what is known as the lack
of dual effect of control: the control actions up to any time t do not affect the state estimation error process for the future
time stages. Using our earlier analysis, it follows then that the optimal control has the form stated in the following:

Theorem 6.2.5 Consider (6.8) with cost criterion given in (6.9). The optimal control is given with

ut = −(B T Pt+1 B + R)−1 B T Pt+1 AE[xt |It ] = −(B T Pt+1 B + R)−1 B T Pt+1 Am̃t ,

with m̃t computed as in (6.21), and Pt generated as in Theorem 5.3.1 with PN = QN . The optimal cost writes as
−1
 NX 
E[m̃T0 P0 m̃0 ] +E w̄kT Pk+1 w̄k + (xk − m̃k ) Q(xk − m̃k ) + E[(xN − m̃N )T QN (xN − m̃N )]
T

k=0

In the above problem, we observed that the optimal control has a separation structure: The controller first estimates the
state, and then applies its control action, by regarding the estimate as the state itself.
Separation of Estimation and Control. In the above we observe that the optimal control policy is the same as that
in the fully observed setup in Theorem 5.3.1, except that the state is replaced with its estimate. The sufficiency of
6.3 On the Controlled Markov Construction in the Space of Probability Measures and Extension to General Spaces 121

conditional expectation in optimal control is generally known as the separation of estimation and control [140] [205]
[334] [218]–that is, the separation principle is said to hold when an optimal control exists in a subset of admissible
policies where the control depends on the information only through the conditional expectation of the state given the
information available–, and for this particular case, a more special version of it, known as the certainty equivalence
principle, applies: As expressed in [29, eqs. (2.20)–(2.22)], a control problem possesses the certainty equivalence (CE)
property if the closed-loop optimal control policy has the same form as the deterministic optimal control policy under
perfect state observation and in the absence of process noise. More precisely, if in the absence of noise the optimal
control policy for the deterministic system is

uk = ϕk (xk ), (6.23)

and CE holds, then the optimal closed-loop control policy for the noisy and not necessarily fully observed system is

uCE
k = ϕk (E[xk |y[0,k] , u[0,k−1] ]), ∀k. (6.24)

As observed above, the absence of dual effect plays a key part in this analysis leading the separation of estimation and
control principle, in taking E[(xt − m̃)T Q(xt − m̃)] out of the optimization over control policies, since it does not
depend on the policy.

Remark 6.5 (Dual Effect vs. Certainty Equivalence). In [30], dual effect is introduced as the property that the moments
of (xt − m̃t ) do not depend on the past applied control actions (leading to a form of probabilistic independence). A
more general condition would be that (xt − m̃t ) does not depend on the past control policies (and not necessarily
the actions) in that the control policies do not alter the realization of the random variable (xt − m̃t ) (see [102] for
an explicit analysis and relaxations). This distinction is important in certain applications in networked control systems
where separation results are particularly important [345] (as probabilistic independence is often too restrictive when
one goes beyond the Gaussian setup). We also note that separation also applies in the linear quadratic setup when the
noise processes are not Gaussian, though of course the conditional estimations will no longer be linear [35, Lemma
5.2.1]. For results involving non-linear measurement models and for a detailed literature review, the reader is referred
to [102].

In many problems, the dual effect of the control is present and, depending on the control policy, the estimation quality
at the controller regarding future states may be affected. As an example, consider a linear system controlled over an
erasure channel, where the controller applies a control, but does not know whether the control reaches the destination
or not. In this case, the control signal which was intended to be sent, does affect the estimation error [175, 286].

6.3 On the Controlled Markov Construction in the Space of Probability Measures and
Extension to General Spaces

In Section 6.1, we observed that we can replace the state with a probability measure valued state. It is important to
provide notions of convergence and continuity on the spaces of probability measures to be able to apply the machinery
of Chapter 5. In view of Theorem 6.1.1, if we can invoke the measurable selection conditions studied earlier (such as
Assumption 5.2.1), we can use the machinery of optimal stochastic control (such as Bellman’s principle) for partially
observed models.
The reader is referred to Appendix D for review of some concepts involving convergences of probability measures.

6.3.1 Non-linear Filter in the Standard Borel setup

The analysis in Section 6.1 applies essentially identically to the standard Borel setup.
122 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

We consider (6.1). Let X be a standard Borel set from which the controlled Markov process {xt , t ∈ Z+ } takes its
values with transition kernel T . Let Y be a standard Borel space, and let the observation channel Q be defined as the
stochastic kernel (regular conditional probability) from X × U to Y such that Q( · |x, u) is a probability measure on the
Borel σ-algebra B(Y) of Y for every (x, u) ∈ X × U and Q(A| · ) : X × U → [0, 1] is a Borel measurable function for
every A ∈ B(Y).
Let a decision maker (DM) be located at the output of an observation channel Q, with inputs xt and outputs yt . An
admissible policy γ is a sequence of control functions {γt , t ∈} such that γt is measurable with respect to the σ-algebra
generated by the information variables

It = {y[0,t] , u[0,t−1] }, t ∈ N, I0 = {y0 },

where
ut = γt (It ), t∈ (6.25)
are the U-valued control action variables. We define ΓA to be the set of all such admissible policies. The joint distribu-
tion of the state, control, and observation processes is determined by (12.1) and the following system dynamics:
Z

P (x0 , y0 ) ∈ B = Q(dy0 |x0 )P0 (dx0 ), B ∈ B(X × Y),
B

where P0 is the prior distribution of the initial state x0 and Q0 is the observation channel, and for t ∈ N
 
P (xt , yt ) ∈ B (x, y, u)[0,t−1] = (x, y, u)[0,t−1]
Z
= Q(dyt |xt )T (dxt |xt−1 , ut−1 ), B ∈ B(X × Y),
B

where T (·|x, u) is a stochastic kernel from X × U to X. This completes the probabilistic description of the partially
observed model. Let a one-stage cost function c : X × U → [0, ∞), which is a Borel measurable function from X × U
to [0, ∞), be given. Then, we denote by J(γ) the cost function of the policy γ ∈ ΓA , which can be, for instance, finite
horizon, discounted cost or average cost criteria. The goal of the control problem is to find an optimal policy γ ∗ that
minimizes J.
As studied earlier, any such problem can be reduced to a completely observable Markov process [352], [262], whose
states are the posterior state distributions or ’beliefs‘ of the observer; that is, the state at time t is

πt ( · ) := P {Xt ∈ · |y0 , . . . , yt , u0 , . . . , ut−1 } ∈ P(X).

We call this equivalent process the filter process . The filter process has state space P(X) and action space U. Recall
again that P(X) is equipped with the Borel σ-algebra generated by the topology of weak convergence, where, under
this topology P(X) is also a standard Borel space.
The transition probability of the filter process can be constructed as follows. As in the countable setup case, we have
the following explicit Bayesian recursion to define F under a mild regularity condition: Let Q be dominated in the
sense that there exists a dominating reference measure λ such that ∀x ∈ X, Q(dy|xn = x) ≪ λ. Then, define the
Radon-Nikodym derivative
dG(yn ∈ ·|xn = x)
g(x, y) = (y)

as the likelihood function (serving as a conditional probability density function) and we can write the filter πn+1
recursively in terms of πn and yn+1 , un explicitly as a Bayesian update:
R
g(xn+1 , yn+1 )T (dxn+1 |xn , un )πn (dxn )
πn+1 (dxn+1 ) =: F (πn , yn+1 , un )(dxn+1 ) = R XR (6.26)
X X
g(xn+1 , yn+1 )T (dxn+1 |xn , un )πn (dxn )
6.3 On the Controlled Markov Construction in the Space of Probability Measures and Extension to General Spaces 123

As earlier in the countable space setup, the transition probability η of the filter process is constructed as follows. If we
define the measurable function F (π, y, u) := P {xt+1 ∈ · |πt = π, ut = u, yt+1 = y} from P(X) × U × Y to P(X)
and use the stochastic kernel P ( · |π, u) = P {yt+1 ∈ · |πt = π, ut = u} from P(X) × U to Y, we can write η as
Z
η( · |π, u) = 1{F (π,y,u)∈ · } P (dy|π, u). (6.27)
Y

In (6.2), we need to show that the expression P (πt+1 ∈ D|πt , ut ) is a regular conditional probability measure; that
is, for every fixed D ∈ B(P(X)), this is a measurable function on P(X) × U and for every πt , ut , it is a conditional
probability measure on P(X). Furthermore, we need to ensure that c̃, the equivalent cost function, is also a measurable
function.
A proof of the first result below can be found in [4] (see Theorem 15.13 in [4] or p. 215 in [53])

Theorem 6.3.1 Let S be a Polish space and M be the set of all measurable and bounded functions f : S → R. Then,
for any f ∈ M , the integral Z
π(dx)f (x)

defines a measurable function on P(S) under the topology of weak convergence.

This is a useful result since it allows us to view integral forms as measurable functions on the space of probability
measures when we work with the topology of weak convergence. The second useful result follows from Theorem 6.3.1
and Theorem 2.1 of Dubins and Freedman [109] and Proposition 7.25 in Bertsekas and Shreve [37].

Theorem 6.3.2 Let S be a Polish space. A function F : P(S) → P(S) is measurable on B(P(S)) (under weak
convergence), if for all B ∈ B(S) (F (·))(B) : P(S) → R is measurable under weak convergence on P(S), that is for
every B ∈ B(S), (F (π))(B) is a measurable function when viewed as a function from P(S) to R.

By Theorem 6.3.2, we have that F is a Borel measurable function, and η is a stochastic kernel.
The above thus establish that under weak convergence topology, (πt , ut ) forms a standard Borel controlled Markov
chain.
As in the countable setup, in the belief-MDP formulation, the one-stage cost function c̃ : P(X) × U → [0, ∞) for the
filter process is given by
Z
c̃(π, u) := c(x, u)π(dx),
X

With cost function c(x, u) continuous and bounded on X) × U, by an application of the generalized dominated conver-
gence
R theorem (see Theorem D.3.1 [211, Theorem 3.5] [287, Theorem 3.5]), we have that that c̃(π, u) = E π [c(x, u)] :=
π(dx)c(x, u) : P(X) × U → R is also continuous and bounded, and thus Borel measurable as a map from P(X) × U
to R.
Hence, the filter process defines a completely observable Markov process with the components (P(X), U, c̃, η).
For the filter process, let us define another information variable sequence as

I˜t = {π[0,t] , u[0,t−1] }, t ∈ N, I˜0 = {π0 }.

Now, building all these together, as in the countable setup, in view of the results in Chapter 5 (notably Theorem
??5.1.1Theorem 5.1.1hen follows that an optimal control policy of the original POMDP will use the belief πt as a
sufficient statistic for optimal policies (see [352], [262]). More precisely, the filter process is equivalent to the original
POMDP in the sense that for any optimal policy using the filter process, one can construct a policy for the original
POMDP which is optimal, or more generally, for any policy which uses I˜t there exists another one which only uses the
filter process πt and which is at least as good as the original policy.
124 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

6.3.2 Continuity Properties of Belief-MDP: Weak Continuity and Wasserstein Continuity of Filter Kernels

Weak Feller Continuity of the Filter Kernel η. Building on [92], [183] and [127], we first study the weak Feller
property of the filter process; that is, the weak Feller property of the kernel defined in (6.3) under two different sets of
assumptions.

Assumption 6.3.1 (i) The transition probability T (·|x, u) is weakly continuous in (x, u), i.e., for any (xn , un ) →
(x, u), T (·|xn , un ) → T (·|x, u) weakly.
(ii) The observation channel Q(·|x, u) is continuous in total variation, i.e., for any (xn , un ) → (x, u), Q(·|xn , un ) →
Q(·|x, u) in total variation.

Assumption 6.3.2 (i) The transition probability T (·|x, u) is continuous in total variation in (x, u), i.e., for any
(xn , un ) → (x, u), T (·|xn , un ) → T (·|x, u) in total variation.
(ii) The observation channel Q(·|x) is independent of the control variable.

Theorem 6.3.3 [127] Under Assumption 6.3.1, the transition probability η(·|z, u), given in (6.27), of the filter process
is weakly continuous in (z, u).

See also [92] for an earlier though slightly more restrictive result along the above.

Theorem 6.3.4 [183] Under Assumption 6.3.2, the transition probability η(·|z, u), given in (6.27), of the filter process
is weakly continuous in (z, u).

A proof for these is given in Section 6.6.


If the cost function c is continuous and bounded, an application of the dominated convergence theorem implies that
c̃(π, u) is also continuous and bounded. If the action set is compact, then under the weak continuity condition noted
above on the non-linear filter, we have that the measurable selection conditions apply, and solutions to the Bellman or
discounted cost optimality equations exist, and accordingly an optimal control policy exists.
See [187, Theorem 7], which builds on [183], for further refinements with explicit moduli of continuity for the weak
Feller property.
We refer the reader also to [127, Theorem 7.1] which establishes weak Feller property under further sets of assumptions.
See [1, 126, 128, 187] for further results on the above weak Feller property.
It should be noted that the some of conditions above cannot be relaxed in general. Notably, if the transition kernel T
is weakly continuous and the measurement kernel Q is weakly continuous, a counterexample noted in [127, Section
4], attributed to Huizhen (Janey) Yu, shows that the kernel η may not be weak Feller. We revise this counterexample
according to our formulation: Let X = {1, 2} × [0, 1] ∋ x = (x1 , x2 ) so that x1 ∈ {1, 2} and x2 ∈ [0, 1]. Let Y = [0, 1]
and T )(dxt+1 |x1t = x1 , x2t = x2 ) = δx1 (dx1t+1 )δx2 (dx2t+1 ), that is the source is a constant source and therefore the
transition kernel is weak Feller. Let the measurements be:

Q(dy|x1 = 1, x2 = a) = δ1 (dy), a ∈ [0, 1]

Q(dy|x1 = 2, x2 = a) = ([0, a]), a ∈ [0, 1]


Observe that Q is also weak Feller. Consider a sequence of x2 values: an = n1 . Let z n = δ[1/2,1/2] (dx1 )δan (dx2 ) and
z = δ[1/2,1/2] (dx1 )δ0 (dx2 ). Now, for every n > 0: we have that the measurement is fully informative with regard to
x1 component of the state:

η(dπt+1 |πt = z n ) = 1/2δ[1,0] (dπt+1 )δan (dx2 ) + 1/2δ[0,1] (dπt+1 )δan (dx2 )

whereas at x2t = 0, we have that the measurement is non-informative with regard to x1 and thus the conditional
probability on x1 does not change:
6.3 On the Controlled Markov Construction in the Space of Probability Measures and Extension to General Spaces 125

η(dπt+1 |πt = z) = δ[1/2,1/2]×δ0 (dπt+1 )

Observe that in the above π n → π weakly. Now consider the continuous and bounded function f (π) = (1 −
d(π, z))1{d(π,z)≤1} where d(π, z) = ∥π(dx1 , [0, 1]) − z(dx1 , [0, 1])∥2 , which is continuous, we observe that
Z Z
1
√ = lim η(dπt+1 |πt = z n )f (πt+1 ) ̸= η(dπt+1 |πt = z)f (πt+1 ) = 0
2 n→∞
Accordingly, we do not have the weak Feller continuity for this example.
On the other hand [348], consider the special case with the state transition kernel T being weakly continuous and where
the measurements satisfy yt = xt , that is, with full state information, in which case the measurement kernel is

Q(dy|x) = δx (dy), (6.28)

and thus the system does not satisfy Assumption 6.3.1 or 6.3.2 (as the measurement kernel is only weakly continuous
but not total variation continuous). In this case, we do have the weak Feller property.
As examples, taken from [183], suppose that the system dynamics and the observation channel are represented as
follows:

xt+1 = H(xt , ut , wt ),
yt = G(xt , ut−1 , vt ),

where wt and vt are i.i.d. noise processes.

(i) Suppose that H(x, u, w) is a continuous function in x and u. Then, the corresponding transition kernel is weakly
continuous. To see this, observe that, for any c ∈ Cb (X), we have
Z Z
c(x1 )T (dx1 |xn0 , un0 ) = c(H(xn0 , un0 , w0 ))µ(dw0 )
Z Z
→ c(H(x0 , u0 , w0 ))µ(dw0 ) = c(x1 )T (dx1 |x0 , u0 ),

where we use µ to denote the probability model of the noise.


(ii) Suppose that G(x, u, v) = g(x, u)+v, where g is a continuous function and Vt admits a continuous density function
φ with respect to some reference measure ν. Then, the channel is continuous in total variation. Notice that under
this setup, we can write Q(dy|x, u) = φ(y − h(x, u))ν(dy). Hence, the density of Q(dy|xn , un ) converges to
the density of Q(dy|x, u) pointwise, and so, Q(dy|xn , un ) converges to Q(dy|x, u) in total variation by Scheffé’s
Lemma [43]. Hence, Q(dy|x, u) is continuous in total variation under these conditions.
(iii)Suppose that we have H(x, u, w) = h(x, u)+w, where f is continuous and wt admits a continuous density function
φ with respect to some reference measure ν. Then, the transition probability is continuous in total variation: with
this setup we have T (dx1 |x0 , u0 ) = φ(x1 −h(x0 , u0 ))ν(dx1 ). Thus, continuity of φ and h guarantees the pointwise
convergence of the densities, so we can conclude that the transition probability is continuous in total variation by
again Scheffé’s Lemma.

Under the above, it follows that for the belief-MDP (P(X), U, c̃, η), Assumption 5.2.1 holds and therefore Theorem
5.2.1 applies: For a finite horizon cost minimization problem, there exists an optimal control policy which is of Markov
type (Markov in the belief state, πt ). The discounted and average cost criteria will be presented in the following section.

Remark 6.6 (Existence results without separation / belief-MDP reduction). Consider a partially observable stochastic
control problem (POMDP) with the following dynamics.

xt+1 = f (xt , ut , wt ), yt = g(xt , vt ).


126 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

If f (·, ·, w) is continuous and g has the form: yt = g(xt ) + vt , with g continuous and wt admitting a continuous density
function η, an existence result can be established building on the measurable selection criteria under weak continuity
in view of Theorem 6.3.3.
Without adopting the belief-MDP reduction method, such an existence result can also be established by a mea-
sure transformation argument and using a strategic measures approach: With η denoting the density of vn , we have
P (yn ∈ B|xn ) = B η(y − g(xn ))dy. With η and g continuous and bounded, taking y n := yn , by writing
R

xn+1 = f (xn , un , wn ) = f (f (xn−1 , un−1 , wn−1 ), un , wn ), and iterating inductively to obtain

xn+1 = hn (x0 , u[0,n−1] , w[0,n−1] ),

for some hn which is continuous in u[0,n−1] for every fixed x0 , w[0,n−1] , one obtains an effective reduced cost (10.31)
that is a continuous function in the control actions. [346, Section 5.4.2] then implies the existence of an optimal control
policy.R This reasoning is also applicable when the measurements are not additive in the noise but with P (yn ∈ B|xn =
x) = B m(y, x)η(dy) for some m continuous in x and η a reference measure.
It may be important to note that Bismut [46] arrived at related results for partially observed models in continuous-time,
through an approach which also avoids separation / the construction of a belief-MDP; see [348] and Section 10.8.1 for
further discussion.

Wasserstein Continuity of the Filter Kernel η. Regularity under the Wasserstein metric has been studied in [1, 99]
and [100]:

Assumption 6.3.3
1. (X, d) is a bounded compact metric space with diameter D (where D = supx,y∈X d(x, y)).
2. The transition probability T (· | x, u) is continuous in total variation in (x, u), i.e., for any (xn , un ) →
(x, u), T (· | xn , un ) → T (· | x, u) in total variation.
3. There exists α ∈ R+ such that
∥T (· | x, u) − T (· | x′ , u)∥T V ≤ αd(x, x′ )
for every x, x′ ∈ X, u ∈ U.
4. There exists K1 ∈ R+ such that
|c(x, u) − c(x′ , u)| ≤ K1 d(x, x′ ).
for every x, x′ ∈ X, u ∈ U.
5. The cost function c is bounded and continuous.

Theorem 6.3.5 [99] Assume that X and Y are Polish spaces. If Assumptions 6.3.3-1,3 are fulfilled, then we have

W1 (η(· | z0 , u), η (· | z0′ , u)) ≤ K2 W1 (z0 , z0′ ) ,


αD(3−2δ(Q))
with K2 := 2 for all z0 , z0′ ∈ P(X), u ∈ U.

Assumption 6.3.4 (i) (X, d) is a compact metric space.


(ii) There exists a constant θ ∈ (0, 1) such that

W1 (T (· | x, u) − T (· | x′ , u)) ≤ θ · d (x, x′ )

for every x, x′ ∈ X, u ∈ U.
(iii)There exists a constant γ ∈ R+ such that

∥Q(· | x) − Q (· | x′ )∥T V ≤ γ · d (x, x′ )


6.3 On the Controlled Markov Construction in the Space of Probability Measures and Extension to General Spaces 127

for every x, x′ ∈ X.

Theorem 6.3.6 [98, Theorem 2.4] Assume that X and Y are Polish spaces. Under Assumption 6.3.4, we have
 
′ 3θγD
W1 (η (· | z0 , u) , η (· | z0 , u)) ≤ θ + W1 (z0 , z0′ )
2

for all z0 , z0′ ∈ Z, u ∈ U, where D = supx,y∈X d(x, y).

Remark 6.7. [348] has presented an alternative approach, without belief-separation, and has arrived further conditions
for the existence of optimal policies for discounted and average cost problems as well as the unique ergodicity property
for both controlled and control-free setups. Such an approach leads to complementary conditions on the weak Feller
property on the state, which considers the entire past as the state endowed with the product topology.

6.3.3 Existence of Optimal Policies: Discounted Cost and Average Cost

Consider the minimization of either the discounted cost criterion (for some β ∈ (0, 1)

X
J(µ, γ) := Eµγ [ β k c(xk , uk )] (6.29)
k=0

or the average cost criterion


N −1
1 γ X
J(µ, γ) := lim sup E [ c(xk , uk )], (6.30)
N →∞ N µ
k=0

over all admissible control policies γ = {γ0 , γ1 , · · · } ∈ Γ with x0 ∼ µ.


Discounted Cost Cost.

Theorem 6.3.7 If the cost function c : X × U → R is continuous and bounded, and U is compact, under under
Theorems 6.3.3 or 6.3.4, for any β ∈ (0, 1), there exists an optimal solution to the discounted cost optimality problem
with a continuous and bounded value function. Furthermore, under Assumption 6.3.3, with K2 = αD(3−2δ(Q)) 2 , if
βK2 < 1 the value function is Lipschitz continuous.

Proof. An application of the dominated convergence theorem implies that c̃(π, u) is also continuous and bounded. If
the action set is compact, then under Theorems 6.3.3 or 6.3.4, which imply that η is weakly continuous, we have that
the measurable selection conditions (see e.g. [164]) apply, and solutions to the Bellman or discounted cost optimality
equations exist, and accordingly an optimal control policy exists. For the second result, Theorem 5.5.3 (see [270, The-
orem 4.37]) leads to Lipschitz regularity under the Wasserstein continuity condition on the kernel. ⋄

Average Cost. The average cost is a significantly more challenging problem as the typical contraction conditions via
minorization (7.2.2) for the kernel η is too demanding. As is studied in detail in Chapter 7, the average cost optimality
equation (ACOE) plays a crucial role for the analysis and the existence results of MDPs under the infinite horizon
average cost optimality criteria. The triplet (h, ρ∗ , γ ∗ ), where h, γ : P(X) → R are measurable functions and ρ∗ ∈ R
is a constant forms the ACOE if
 Z 

h(z) + ρ = inf c̃(z, u) + h(z1 )η(dz1 |z, u)
u∈U
Z
= c̃(z, γ ∗ (z)) + h(z1 )η(dz1 |z, γ ∗ (z)) (6.31)
128 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

for all z ∈ P(X). It is well known that (see e.g. [164, Theorem 5.2.4]) if (6.31) is satisfied with the triplet (h, ρ∗ , γ ∗ ),
and furthermore if h satisfies
Ezγ [h(Zt )]
sup lim = 0, ∀z ∈ P(X)
γ∈Γ t→∞ t

then γ ∗ is an optimal policy for the POMDP under the infinite horizon average cost optimality criteria, and

J ∗ (z) = inf J(z, γ) = ρ∗ ∀z ∈ P(X).


γ∈Γ

Theorem 6.3.8 [99] Under Assumption 6.3.3, with K2 = αD(3−2δ(Q)) 2 < 1, a solution to the average cost optimality
equation (ACOE) exists. This leads to the existence of an optimal control policy, and optimal cost is constant for every
initial state.

The proof follows from Corollary 7.3.1. For belief-MDPs, we should emphasize that minorization conditions (as in
Assumption 7.2.2) are typically not applicable.

6.3.4 A useful structural result: Concavity of the value function in the priors

The following theorem establishes concavity of the optimal cost in a single-stage stochastic control problem over the
space of initial distributions and this also applies for multi-stage setups.
R
Theorem 6.3.9 Let c(x, γ(y))P Q(dx, dy) exist for all γ ∈ Γ and P ∈ P(X). Then,

J ∗ (P, Q) = inf EPQ,γ [c(x, γ(y))]


γ∈Γ

is concave in P .

Proof. For a ∈ [0, 1] and P ′ , P ′′ ∈ P(X) we let P = aP ′ + (1 − a)P ′′ . Note that P Q = aP ′ Q + (1 − a)P ′′ Q. We
have

J(aP ′ + (1 − a)P ′′ , Q) = J(P, Q)


= inf EPQ,γ [c(x, γ(y))]
γ∈Γ
Z
= inf c(x, γ(y))P Q(dx, dy)
γ∈Γ
 Z
= inf a c(x, γ(y))P ′ Q(dx, dy)
γ∈Γ
Z 
+(1 − a) c(x, γ(y))P ′′ Q(dx, dy)
 Z 
≥ inf a c(x, γ(y))P ′ Q(dx, dy)
γ∈Γ
 Z 
+ inf (1 − a) c(x, γ(y))P ′′ Q(dx, dy)
γ∈Γ

= aJ(P ′ , Q) + (1 − a)J(P ′′ , Q)


6.4 Filter Stability 129

6.4 Filter Stability

The filter stability problem refers to the correction of an incorrectly initialized non-linear filter for a partially observed
stochastic dynamical system (controlled or control-free) with increasing measurements. Let us describe this property
more explicitly: Given a prior µ ∈ P(X) and a policy γ ∈ Γ we can then define the filter and predictor for a POMDP
using the (strategic) measure P µ,γ .

Definition 6.4.1 (i) We define the one step predictor process as the sequence of conditional probability measures
µ,γ
πn− (·) = P µ,γ (Xn ∈ ·|Y[0,n−1] , U[0,n−1] ) = P µ,γ (Xn ∈ ·|Y[0,n−1] ) n ∈ N

(ii) We define the filter process as the sequence of conditional probability measures

πnµ,γ (·) = P µ,γ (Xn ∈ ·|Y[0,n] , U[0,n−1] ) = P µ,γ (Xn ∈ ·|Y[0,n] ), n ∈ Z+ (6.32)

Remark 6.8. Recall that the U[0,n−1] are all functions of the Y[0,n−1] , so conditioning on the control actions is not
necessary in the above definitions. Yet this conditional probability would be policy dependent; if we condition on the
past actions, this conditioning would be policy-independent.

Say a prior µ ∈ P(X) and a policy γ ∈ Γ are chosen, an observer sees measurements Y[0,∞) generated via the
strategic measure P µ,γ . The observer is aware that the policy applied is γ, but incorrectly thinks the prior is ν ̸= µ. The
observer will then compute the incorrectly initialized filter πnν,γ while the true filter is πnµ,γ . The filter stability problem
is concerned with the merging of πnν,γ and πnµ,γ as n goes to infinity.
It will be useful to note that the filter is the strategic measure conditioned on the sigma field F0,n
Y
and restricted to the
sigma field Fn .
X

πnµ,γ (·) = P µ,γ (Xn ∈ ·|Y[0,n] ) = P µ,γ |FnX |F0,n


Y

In the literature, there are a number of merging notions when one considers stability which we enumerate here. Let
Cb (X) represent the set of continuous and bounded functions from X → R.
R R
Definition 6.4.2 Two sequences of probability measures Pn , Qn merge weakly if ∀ f ∈ Cb (X) we have limn→∞ f dPn − f dQn =
0.

Definition 6.4.3
R For two R probability measures P and Q we define the total variation norm as ∥P − Q∥T V =
sup∥f ∥∞ ≤1 f dP − f dQ where f is assumed measurable. We say two sequences of probability measures Pn ,
Qn merge in total variation if ∥Pn − Qn ∥T V → 0 as n → ∞.

Definition 6.4.4
dP dP dP
R R
(i) For two probability measures P and Q we define the relative entropy as D(P ∥Q) = log dQ dP = dQ log dQ dQ
dP
where we assume P ≪ Q and dQ denotes the Radon-Nikodym derivative of P with respect to Q.

(ii) Let X and Y be two random variables, let P and Q be two different joint measures for (X, Y ) with P ≪ Q. Then
we define the (conditional) relative entropy between P (X|Y ) and Q(X|Y ) as

dPX|Y
Z  
D(P (X|Y )∥Q(X|Y )) = log (x, y) dP (x, y)
dQX|Y
dPX|Y
Z Z   
= log (x, y) dP (x|Y = y) dP (y) (6.33)
dQX|Y

We define here the different notions of stability for the filter:


130 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

Definition 6.4.5 (i) A filter process is said to be stable in the sense of weak merging with respect to a policy γ P µ,γ
almost surely (a.s.) if there exists a set of measurement sequences A ⊂ Y Z+ with P µ,γ probability 1 such that for
any sequence in A;Rfor any f ∈ RCb (X ) and any prior ν with µ ≪ ν (i.e., for all Borel B ν(B) = 0 =⇒ µ(B) = 0)
we have limn→∞ f dπnµ,γ − f dπnν,γ = 0.
(ii) A filter process is said to be stable in the sense of total variation in expectation with respect to a policy γ if for any
measure ν with µ ≪ ν we have limn→∞ E µ,γ [∥πnµ,γ − πnν,γ ∥T V ] = 0.
(iii)A filter process is said to be stable in the sense of total variation with respect to a policy γ P µ,γ a.s. if there exists a
set of measurement sequences A ⊂ Y Z+ with P µ,γ probability 1 such that for any sequence in A; for any measure
ν with µ ≪ ν we have limn→∞ ∥πnµ,γ − πnν,γ ∥T V = 0 P µ,γ a.s..
(iv) A filter process is said to be stable in the sense of relative entropy with respect to a policy γ if for any measure ν
with µ ≪ ν we have limn→∞ E µ,γ [D(πnµ,γ ∥πnν,γ )] = 0.
(v) The filter is said to be universally stable in one of the above notions if the notion holds with respect to every
admissible policy γ ∈ Γ .

Predictor stability is defined in an analogous fashion for each of the criteria above.
Total variation merging implies weak merging, and relative entropy merging (i.e. D(Pn ∥Qn ) → 0) implies total
variation merging via Pinsker’s inequality [93].
One of the main differences between control-free and controlled partially observed Markov chains is that the filter
is always Markovian under the former, whereas under a controlled model the filter process may not be Markovian
since the control policy may depend on past measurements in an arbitrary (measurable) fashion. This complicates the
dependency structure and therefore results from the control-free case do not directly apply to the controlled setup.
We made the observation earlier that under observability and a controllability assumption, any incorrectly initialized fil-
ter will converge to the correct Kalman filter (we note that partial convergence and robustness results on the asymptotic
equivalence of conditional expectations and linear estimates for non-Gaussian priors for linear systems are reported
in [292]). In the following, we will present a concise discussion on how such results carry over to the stochastic non-
linear setup.
Much of the results on filter stability involves control-free systems. Thus, results have considered partially observed
Markov processes (POMP) as opposed to partially observed Markov decision processes (POMDP). Since there is no
control in such systems, there is no past dependency in the system and the pair (Xn , Yn )∞ n=0 is always a Markov
chain. For such control-free models, filter stability has been studied extensively and we refer the reader to [85] for
a comprehensive review and a collection of different approaches. As discussed in [85], filter stability arises via two
separate mechanisms:
1. The transition kernel is in some sense sufficiently ergodic, forgetting the initial measure and therefore passing this
insensitivity (to incorrect initializations) on to the filter process.
2. The measurement channel provides sufficient information about the underlying state, allowing the filter to track the
true state process.
To be able to present a concise discussion, building on some prior material in the notes, for both controlled and control-
free setups we review conditions in [224] based on Dobrushin’s coefficients of the measurement channel and the
controlled transition kernel. Recall (3.29). We consider a slight generalization in the following.

Definition 6.4.6 [105, Equation 1.16] For a kernel operator K : S1 → P(S2 ) (that is a regular conditional probability
from S1 to S2 ) for standard Borel spaces S1 , S2 , we define the Dobrushin coefficient as:
n
X
δ(K) = inf min(K(x, Ai ), K(y, Ai )) (6.34)
i=1

where the infimum is over all x, y ∈ S1 and all partitions {Ai }ni=1 of S2 .
6.4 Filter Stability 131

Let us define
δ̃(T ) := inf δ(T (·|·, u)).
u∈U

The following can be viewed as a generalization of Theorem 3.1.8.

Theorem 6.4.1 [224, Theorem 3.3] Assume that for µ, ν ∈ P(X), we have µ ≪ ν. Then we have
 µ,γ ν,γ
E µ,γ ∥πn+1 ∥T V ≤ (1 − δ̃(T ))(2 − δ(Q))E µ,γ [∥πnµ,γ − πnν,γ ∥T V ] .

− πn+1

In particular, defining α := (1 − δ̃(T ))(2 − δ(Q)), we have

E µ,γ [∥πnµ,γ − πnν,γ ∥T V ] ≤ 2αn .

By applying the Borel-Cantelli lemma and Markov’s inequality, we have that exponential stability in expectation im-
plies the same result in an almost sure sense as well: assume that the filter is exponentially stable with coefficient
α = (1 − δ(T ))(2 − δ(Q) < 1 and let ρ be a value ρ < α1 . Then we have for every ϵ > 0,
∞ ∞
X X E µ [∥πnµ − πnν ∥T V ]
P µ (ρk ∥πnµ − πnν ∥T V ≥ ϵ) ≤ ρk
ϵ
k=0 k=0

∥µ − ν∥T V X
≤ (ρα)k
ϵ
k=0
∥µ − ν∥T V 1
=
ϵ 1 − ρα
<∞
1
thus by Borel Cantelli Lemma ρk ∥πnµ − πnν ∥T V → 0 P µ a.s. for any ρ < α. See [224, Remark 3.10].
This also establishes that the rate of convergence is uniform over all priors ν as long as µ ≪ ν.
A further approach to ensuring filter stability, via sample paths, is by an analysis which builds on the Hilbert projective
metric: [145]

Definition 6.4.7 Two non-negative measures µ, ν on (X, B(X)) are comparable, if there exist positive constants 0 <
a ≤ b, such that
aν(A) ≤ µ(A) ≤ bν(A)
for any Borel subset A ⊂ X.

Definition 6.4.8 (Mixing kernel) The non-negative kernel K defined on X is mixing, if there exists a constant 0 < ε ≤
1, and a non-negative measure λ on X, such that
1
ελ(A) ≤ K(x, A) ≤ λ(A)
ε
for any x ∈ X, and any Borel subset A ⊂ X.

Definition 6.4.9 (Hilbert metric). Let µ, ν be two non-negative finite measures. We define the Hilbert metric on such
measures as   
µ(A)
supA|ν(A)>0 ν(A)

 log µ(A) if µ, ν are comparable
 inf A|ν(A)>0 ν(A)
h(µ, ν) = (6.35)

 0 if µ = ν = 0


else
132 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

Note that h(aµ, bν) = h(µ, ν) for any positive scalars a, b. Therefore, the Hilbert metric is a useful metric for nonlinear
filters since it is invariant under normalization, and the following lemma demonstrates that it bounds the total-variation
distance.

Lemma 6.4.1 [145, Lemma 3.4] Let µ, ν be two non-negative finite measures,
2
i. ∥µ − ν∥T V ≤ log 3 h (µ, ν) .

ii. If the nonnegative kernel K is a mixing kernel (see Definition 6.4.8) with constant ϵ, then h (Kµ, Kν) ≤
1
ε2 ∥µ − ν∥T V .

Lemma 6.4.2 ( [145], Lemma 3.8) The nonnegative linear operator τ on M+ (X) (positive measures on X) associated
with a nonnegative kernel K defined on X
 
h (Kµ, Kν) 1
τ (K) := sup = tanh H(K)
0<h(µ,ν)<∞ h (µ, ν) 4

where
H(K) := sup h (Kµ, Kν)
µ,ν

is over nonnegative measures, is a contraction (called the Birkhoff contraction coefficient) , is a contraction under the
Hilbert metric if H(K) < ∞ (which implies τ (K) < 1).

A controlled version of a contraction via the Hilbert metric [145] can be obtained [100]:
Recall that
F (z, y, u)(·) = P {xk+1 ∈ · | πk = z, Yk+1 = y, Uk = u}

Assumption 6.4.1 1. Q(y|x) ≥ ϵ > 0 for every x ∈ X and y ∈ Y.


2. The transition kernel T (.|., u) is a mixing kernel (see Definition 6.4.8) for every u ∈ U.

Lemma 6.4.3 [1] Under Assumption 8.3.1, there exists a constant r < 1 such that

h(F (µ, y, u), F (ν, y, u)) ≤ rh(µ, ν) (6.36)

1−ϵ2u ϵ
for every comparable µ, ν ∈ P(X) and for every u ∈ U and y ∈ Y. Here r = 1+ϵ2u ϵ , ϵu is the mixing constant of the
kernel T (.|., u).

Another filter stability result which will also be useful in numerical methods for POMDPs to be considered later is via
the following stochastic non-linear observability definition.

Definition 6.4.10 [Stochastic Observability for Non-Linear Systems] [225] A POMDP is called one step observable
(universal in admissible control policies) if for every f ∈ Cb (X) and every ϵ > 0 there exists a measurable and
bounded function g such that
Z
∥f (·) − g(y)Q(dy|·)∥∞ < ϵ (6.37)
Y

Theorem 6.4.2 [225] Assume that µ ≪ ν and that the POMDP is one step observable. Then the predictor is univer-
sally stable weakly a.s. .

We now present an example for observability.


6.4 Filter Stability 133

Example 6.9. [226] Consider a finite setup X = {a1 , · · · , an } and let the noise space be V = {b1 , · · · , bm }. Now,
assume y = h(x, v) has K distinct outputs, where 1 ≤ K ≤ (n)(m) and Y = {c1 , · · · , cK }. We note that for such
a setup, there is already a sufficient and necessary condition for filter stability provided in [315, Theorem V.2] (see
also [313]. We examine this case to show that Definition 6.38 above leads to filter stability.
For each x, hx can be viewed as a partition of V, assigning each bi ∈ X to an output level cj ∈ Y. We can track this by
the matrix Hx (i, j) = 1 if hx (bi ) = yj and zero else. Let Q be the 1 × m vector representing the probability measure
  this can be generalized for the control-free case to N -step
of the noise. We consider the one step observability (though
α1
 α2  R
observability for N > 1). Let g(ci ) = αi , with α =  .  and V g(h(x, v))Q(dv) =: (QHx )α. Any function f (x)
 
.
 . 
αK
can be expressed as a n × 1 vector and hence the question reducesto finding a vector α so that f = QHα, and the
QHa1
 .. 
system is one step observable if and only if the matrix A ≡  .  is rank n.
QHan
Since the spaces are finite, Theorem 6.4.4, to be presented below, leads to filter stability, both in total variation and
weakly, in expectation.

Further examples for measurement channels satisfying Definition 6.4.10 have been reported in [226, Section 3]2 .
The observability notion defined above only results in stability of the predictor in the weak sense P µ,γ a.s. .We next
extend this stability to total variation P µ,γ a.s. .Let the measurement channel Q be dominated in the sense that there
exists a reference measure λ such that ∀x ∈ X , Q(Y ∈ ·|xn = x) ≪ λ(·). Then, we define the Radon-Nikodym
derivative
dQ(Yn ∈ ·|xn = x)
q(x, y) := (y) (6.39)

which serves as a likelihood function. We will consider one of the following assumptions.

Assumption 6.4.2 (i) T (·|x, u) is absolutely continuous with respect to a dominating measure ϕ for every x ∈ X , u ∈
U, so that t(x1 , x, u) = dT (·|x,u)
dϕ (x1 ) where t is continuous in x for every x1 ∈ X and u ∈ U.

(ii) q(x, y) is bounded and continuous in x for every fixed y. Furthermore, q(x, y) > 0 for all x ∈ X , y ∈ Y.

Assumption 6.4.3 T (·|x, u) is absolutely continuous with respect to a dominating measure ϕ for every x ∈ X , u ∈ U,
so that s(x1 , x, u) = dT (·|x,u)
dϕ (x1 ). The family of (conditional densities) {s(·, x, u)}x∈X ,u∈U is uniformly bounded
and equicontinuous.

Theorem 6.4.3 [225] Let µ ≪ ν. Let Assumption 6.4.2 or Assumption 6.4.3 hold. If the predictor is universally stable
in the weak sense a.s. then it is also universally stable in total variation a.s. .

One of the key steps in the proof of Theorem 6.4.2 is that P µ,γ (Yn ∈ ·|Y[0,n−1] ) and P ν,γ (Yn ∈ ·|Y[0,n−1] ) merge
in total variation P µ,γ a.s. as n → ∞. To achieve this in a POMDP, we apply Blackwell and Dubins [50] to the
2
For control-free systems, [226] defines the following: A control-free filter is N -step observable if for every f ∈ Cb (X ) and every
ϵ > 0 there exists a measurable and bounded function g such that
Z
∥f (·) − g(y[1,N ] )Q(dy[1,N ] |X1 = ·)∥∞ < ϵ (6.38)
Y

A further notion is observability: A POMP is observable if for every f ∈ Cb (X ) and every ϵ > 0 there exists N ∈ N and a measurable
and bounded function g such that (6.38) applies. Due to the presence of dual effect, these N -step definitions require a more refined
approach for filter stability.
134 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

measurement process {Yn }∞ n=0 . However, [50] is fundamentally about predictive measures of the future given the past,
and hence only directly implies predictor stability results, not the filter. Filter stability is studied next.

Assumption 6.4.4 The measurement channel Q is continuous in total variation. That is, for any sequence an → a ∈ X
we have ∥Q(·|an ) − Q(·|a)∥T V → 0 or in other words ∥P (Y0 ∈ ·|X0 = an ) − P (Y0 ∈ ·|X0 = a)∥T V → 0.

Assumption 6.4.2(ii), together with the related domination condition (6.39), implies Assumption 6.4.4 (see [183, Sec-
tion 2.3]); see also [183, Theorem 3] for a partial converse result.

Theorem 6.4.4 [225]


(i) Let Assumption 6.4.4 hold. If the predictor is universally stable in weak merging a.s. , then the filter is universally
stable in weak merging in expectation.
(ii) The filter is universally stable in total variation in expectation if and only if the predictor is universally stable in
total variation in expectation.
(iii)The filter is universally stable in total variation in expectation if and only if it is universally stable in total variation
a.s. .
(iv) Let µ ≪ ν, and assume for any policy γ there exists some finite n such that E µ,γ [D(πnµ,γ ∥πnµ,γ )] < ∞ and some
m such that E µ,γ [D(P µ,γ |F Y ∥(P ν,γ |F Y )] < ∞. Then the filter is universally stable in relative entropy if and
0,m 0,m
only if it is universally stable in total variation in expectation.

Applications of these will be discussed in the context of numerical methods for POMDPs later in the notes. Filter
stability is also related to robustness of optimal costs to incorrect initializations for controlled models [225].

6.5 Bibliographic Notes

Earlier work on separation results for partially observed Markov Decision Processes include [352], [296], [262]. For
linear systems, classical texts include [10, 11, 35, 73, 199, 205, 206]. See [85] for a comprehensive review on filter
stability. A very comprehensive recent book on POMDPs is [196].
It has been shown relatively recently that one could approach the Riccati/Kalman Filter updates as a contraction map in
positive-definite matrices [68] (see also [214] and [216]), leading to a concise and direct proof of convergence as well
as stability (though with strict controllability and observability conditions, instead of detectability and stabilizability).
On filter stability, related work in the control-free domain includes [84, 85, 158].

6.6 Appendix

6.6.1 Proof of Theorems 6.3.4 and 6.3.3.

We present the unified proof given in [183]. We first recall (D.3) which is used to metrize weak convergence. The following result
plays a key role.

Lemma 6.6.1 [183] Let X be a Borel space. Suppose that we have a family of uniformly bounded real Borel measurable functions
{fn,λ }n≥1,λ∈Λ and {fλ }λ∈Λ , for some set Λ. If, for any xn → x in X, we have

lim sup |fn,λ (xn ) − fλ (x)| = 0 (6.40)


n→∞ λ∈Λ
6.6 Appendix 135

lim sup |fλ (xn ) − fλ (x)| = 0, (6.41)


n→∞ λ∈Λ

then, for any µn → µ weakly in P(X), we have


Z Z
lim sup fn,λ (x)µn (dx) − fλ (x)µ(dx) = 0.
n→∞ λ∈Λ X X

In Theorem 6.3.3 and Theorem 6.3.4, we need to show that, for every (z0n , un ) → (z0 , u) in Z × U, we have
Z Z
sup f (z1 )η(dz1 |z0n , un ) − f (z1 )η(dz1 |z0 , u) → 0,
∥f ∥BL ≤1 Z Z

where we equip Z with the metric ρ to define bounded-Lipschitz norm ∥f ∥BL of any Borel measurable function f : Z → R. We
can equivalently write this as
Z Z
sup f (z1 (z0n , un , y1 ))P (dy1 |z0n , un ) − f (z1 (z0 , u, y1 ))P (dy1 |z0 , u) → 0. (6.42)
∥f ∥BL ≤1 Y Y

The term in equation (6.42) can be upper bounded as follows:


Z Z
sup f (z1 (z0n , un , y1 ))P (dy1 |z0n , un ) − f (z1 (z0 , u, y1 ))P (dy1 |z0 , u)
∥f ∥BL ≤1 Y Y
Z Z
≤ sup f (z1 (z0n , un , y1 ))P (dy1 |z0n , un ) − f (z1 (z0n , un , y1 ))P (dy1 |z0 , u)
∥f ∥BL ≤1 Y Y
Z
+ sup f (z1 (z0n , un , y1 )) − f (z1 (z0 , u, y1 )) P (dy1 |z0 , u)
∥f ∥BL ≤1 Y

≤ ∥P (·|z0n , un ) − P (·|z0 , u)∥T V


Z
+ sup f (z1 (z0n , un , y1 )) − f (z1 (z0 , u, y1 )) P (dy1 |z0 , u), (6.43)
∥f ∥BL ≤1 Y

where, in the last inequality, we have used ∥f ∥∞ ≤ ∥f ∥BL ≤ 1. To prove that (6.43) (and so (6.42)) goes to 0, it is sufficient to
establish the following results:

(i) P (dy1 |z0 , u0 ) is continuous in total variation,

(ii) limn→∞ Y ρ(z1 (z0n , un , y1 ), z1 (z0 , u, y1 ))P (dy1 |z0 , u) = 0 as (z0n , un ) → (z0 , u).
R

Indeed, suppose that (i) and (ii) hold. Then, the first term in (6.43) goes to 0 as P (·|z0 , u) is continuous in total variation. For the
second term in (6.43), we have
Z
sup f (z1 (z0n , un , y1 )) − f (z1 (z0 , u, y1 )) P (dy1 |z0 , u)
∥f ∥BL ≤1 Y
Z
≤ ρ(z1 (z0n , un , y1 ), z1 (z0 , u, y1 ))P (dy1 |z0 , u)
Y
→ 0 as n → ∞ (by (ii)).

Therefore, to complete the proof of Theorem 6.3.3 and Theorem 6.3.4, we will prove (i) and (ii).

Proof of Theorem 6.3.3

We first prove (i); that is, P (dy1 |z0 , u) is continuous in total variation. To this end, let (z0n , un ) → (z0 , u). Then, we write

sup P (A|z0n , un ) − P (A|z0 , u)


A∈B(Y)
Z Z
= sup Q(A|x1 , un )T (dx1 |z0n , un ) − Q(A|x1 , u)T (dx1 |z0 , u) ,
A∈B(Y) X X
136 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

where T (dx1 |z0n , un ) := X T (dx1 |x0 , un )z0n (dx0 ). Note that, n


R
R by Lemma 6.6.1, we can show that T (dx
R 1 |z0 , un ) → T (dx1 |z0 , u)
weakly. Indeed, if g ∈ Cb (X), then we define rn (x0 ) = X g(x1 )T (dx1 |x0 , un ) and r(x0 ) = X g(x1 )T (dx1 |x0 , u). Since
T (dx1 |x0 , u) is weakly continuous, we have rn (xn n
0 ) → r(x0 ) when x0 → x0 . Hence, by Lemma 6.6.1, we have
Z Z
lim rn (x0 )z0n (dx0 ) − r(x0 )z0 (dx0 ) = 0.
n→∞ X X

Hence, T (dx1 |z0n , un )


→ T (dx1 |z0 , u) weakly. Moreover, the families of functions {Q(A| · , un )}n≥1,A∈B(Y) and {Q(A| · , u)}A∈B(Y)
satisfy the conditions of Lemma 6.6.1 as Q is continuous in total variation distance. Therefore, Lemma 6.6.1 yields that
Z Z
lim sup Q(A|x1 , un )T (dx1 |z0n , un ) − Q(A|x1 , u)T (dx1 |z0 , u) = 0.
n→∞ A∈B(Y) X X

Thus, P (dy1 |z0 , u) is continuous in total variation.

To prove (ii), we write


Z
ρ(z1 (z0n , un , y1 ), z1 (z0 , u, y1 ))P (dy1 |z0 , u)
Y

Z X Z
= 2−m+1 fm (x1 )z1 (z0n , un , y1 )(dx1 )
Y m=1 X
Z
− fm (x1 )z1 (z0 , u, y1 )(dx1 ) P (dy1 |z0 , u)
X

X Z Z
= 2−m+1 fm (x1 )z1 (z0n , un , y1 )(dx1 )
m=1 Y X
Z
− fm (x1 )z1 (z0 , u, y1 )(dx1 ) P (dy1 |z0 , u),
X

where we have used Fubini’s theorem with the fact that supm ∥fm ∥∞ ≤ 1. For each m, let us define
 Z Z 
(n)
I+ := y1 ∈ Y : fm (x1 )z1 (z0n , un , y1 )(dx1 ) > fm (x1 )z1 (z0 , u, y1 )(dx1 )
X X
 Z Z 
(n) n
I− := y1 ∈ Y : fm (x1 )z1 (z0 , un , y1 )(dx1 ) ≤ fm (x1 )z1 (z0 , u, y1 )(dx1 ) . (6.44)
X X

Then, we can write


Z Z Z
fm (x1 )z1 (z0n , un , y1 )(dx1 ) − fm (x1 )z1 (z0 , u, y1 )(dx1 ) P (dy1 |z0 , u)
Y X X
Z Z Z 
= fm (x1 )z1 (z0n , un , y1 )(dx1 ) − fm (x1 )z1 (z0 , u, y1 )(dx1 ) P (dy1 |z0 , u)
(n)
I+ X X
Z Z Z 
+ fm (x1 )z1 (z0 , u, y1 )(dx1 ) − fm (x1 )z1 (z0n , un , y1 )(dx1 ) P (dy1 |z0 , u).
(n)
I− X X

(n)
In the sequel, we only consider the term with the set I+ . The analysis for the other one follows from the same steps. We have
Z Z Z 
fm (x1 )z1 (z0n , un , y1 )(dx1 ) − fm (x1 )z1 (z0 , u, y1 )(dx1 ) P (dy1 |z0 , u)
(n)
I+ X X
Z Z
≤ fm (x1 )z1 (z0n , un , y1 )(dx1 )P (dy1 |z0 , u)
(n)
I+ X
Z Z
− fm (x1 )z1 (z0n , un , y1 )(dx1 )P (dy1 |z0n , un )
(n)
I+ X
Z Z
+ fm (x1 )z1 (z0n , un , y1 )(dx1 )P (dy1 |z0n , un )
(n)
I+ X
Z Z
− fm (x1 )z1 (z0 , u, y1 )(dx1 )P (dy1 |z0 , u)
(n)
I+ X
6.6 Appendix 137

≤ ∥P (dy1 |z0 , u) − P (dy1 |z0n , un )∥T V


Z Z
+ fm (x1 )Q(dy1 |x1 , un )T (dx1 |z0n , un )
(n)
X I+
Z Z
− fm (x1 )Q(dy1 |x1 , u)T (dx1 |z0 , u),
(n)
X I+

where we have used ∥fm ∥∞ ≤ 1 in the last inequality. The first term above goes to 0 since P (dy1 |z0 , u) is continuous in total
variation. For the second term, we use Lemma 6.6.1. Indeed, families of functions {fm (·)Q(A| · , un ) : n ≥ 1, A ∈ B(Y)} and
{fm (·)Q(A| · , u) : A ∈ B(Y)} satisfy the conditions in Lemma 6.6.1 as Q is continuous in total variation. Hence, the second term
converges to 0 by Lemma 6.6.1 since T (dx1 |z0n , un ) → T (dx1 |z0 , u) weakly. Hence, for each m, we have
Z Z
lim fm (x1 )z1 (z0n , un , y1 )(dx1 )
n→∞ Y X
Z
− fm (x1 )z1 (z0 , u, y1 )(dx1 ) P (dy1 |z0 , u) = 0.
X

By the dominated convergence theorem, we then have


Z
lim ρ(z1 (z0n , un , y1 ), z1 (z0 , u, y1 ))P (dy1 |z0 , u)
n→∞ Y

X Z Z
≤ 2−m+1 lim fm (x1 )z1 (z0n , un , y1 )(dx1 )
n→∞ Y X
m=1
Z
− fm (x1 )z1 (z0 , u, y1 )(dx1 ) P (dy1 |z0 , u) = 0.
X

This establishes (ii), which completes the proof together with (i).

Proof of Theorem 6.3.4

We first show (i); that is, P (dy1 |z0 , u0 ) is continuous total variation. Let (z0n , un ) → (z0 , u). Then, we have

sup |P (A|z0n , un ) − P (A|z0 , u)|


A∈B(Y)
Z Z
= sup Q(A|x1 )T (dx1 |x0 , un )z0n (dx0 )
A∈B(Y) X X
Z Z
− Q(A|x1 )T (dx1 |x0 , u)z0 (dx0 ) .
X X

For each A ∈ B(Y) and n ≥ 1, we define


Z
fn,A (x0 ) = Q(A|x1 )T (dx1 |x0 , un )
X

and Z
fA (x0 ) = Q(A|x1 )T (dx1 |x0 , u).
X
Then, for all xn
0 → x0 , we have

lim sup |fn,A (xn


0 ) − fA (x0 )|
n→∞ A∈B(Y)
Z Z
= lim sup Q(A|x1 )T (dx1 |xn
0 , un ) − Q(A|x1 )T (dx1 |x0 , u)
n→∞ A∈B(Y) X X

≤ lim ∥T (dx1 |xn


0 , un ) − T (dx1 |x0 , u)∥T V = 0
n→∞

and

lim sup |fA (xn


0 ) − fA (x0 )|
n→∞ A∈B(Y)
138 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter
Z Z
= lim sup Q(A|x1 )T (dx1 |xn
0 , u) − Q(A|x1 )T (dx1 |x0 , u)
n→∞ A∈B(Y) X X

≤ lim ∥T (dx1 |xn


0 , u) − T (dx1 |x0 , u)∥T V = 0.
n→∞

Then, by Lemma 6.6.1, we have


Z Z
lim sup fn,A (x0 )z0n (dx0 ) − fA (x0 )z0 (dx0 )
n→∞ A∈B(Y) X X
Z Z
= lim sup Q(A|x1 )T (dx1 |x0 , un )z0n (dx0 )
n→∞ A∈B(Y) X X
Z Z
− Q(A|x1 )T (dx1 |x0 , u)z0 (dx0 )
X X
= 0.

Hence, P (dy1 |z0 , u0 ) is continuous in total variation.

Now, we show (ii); that is, for any (z0n , un ) → (z0 , u), we have
Z
lim ρ(z1 (z0n , un , y1 ), z1 (z0 , u, y1 ))P (dy1 |z0 , u) = 0.
n→∞ Y

From the proof of Theorem 6.3.3, it suffices to show that


Z Z
lim fm (x1 )Q(dy1 |x1 )T (dx1 |z0n , un )
n→∞ X I+
(n)

Z Z
− fm (x1 )Q(dy1 |x1 )T (dx1 |z0 , u) = 0. (6.45)
(n)
X I+

Indeed, we have
Z Z Z Z
fm (x1 )Q(dy1 |x1 )T (dx1 |z0n , un ) − fm (x1 )Q(dy1 |x1 )T (dx1 |z0 , u)
(n) (n)
X I+ X I+
Z
(n)
≤ fm (x1 )Q(I+ |x1 )T (dx1 |x0 , un )z0n (dx0 )
X2
Z
(n)
− fm (x1 )Q(I+ |x1 )T (dx1 |x0 , u)z0n (dx0 )
X2
Z
(n)
+ fm (x1 )Q(I+ |x1 )T (dx1 |x0 , u)z0n (dx0 )
X2
Z
(n)
− fm (x1 )Q(I+ |x1 )T (dx1 |x0 , u)z0 (dx0 )
X2
Z
≤ ∥T (dx1 |x0 , un ) − T (dx1 |x0 , u)∥T V z0n (dx0 )
X
Z
(n)
+ fm (x1 )Q(I+ |x1 )T (dx1 |x0 , u)z0n (dx0 )
X2
Z
(n)
− fm (x1 )Q(I+ |x1 )T (dx1 |x0 , u)z0 (dx0 ) ,
X2

(n)
where we have used supn≥1 supx1 ∈X fm (x1 )Q(I+ |x1 ) ≤ 1 in the last inequality. If we define rn (x0 ) = ∥T (dx1 |x0 , un ) −
T (dx1 |x0 , u)∥T V , then rn (xn n
0 ) → 0 whenever x0 → x0 . Then, the first term converges to 0 by Lemma 6.6.1 as z0 → z0
n
R (n)
weakly. The second term also converges to 0 by Lemma 6.6.1, since { X f (x1 )Q(I+ |x1 )T (dx1 |·, u) : n ≥ 1} is a family of
uniformly bounded and equicontinuous functions by total variation continuity of T (dx1 |x0 , u). This proves (ii) and completes the
proof together with (i).


6.7 Exercises 139

6.7 Exercises

Exercise 6.7.1 Consider a linear system with the following dynamics:

xt+1 = axt + ut + wt ,

and let the controller have access to the observations given by:

yt = pt (xt + vt ).

Here {wt , vt , t ∈ Z} are independent, zero-mean, Gaussian random variables, with variances E[w2 ] and E[v 2 ]. The
controller at time t ∈ Z has access to It = {ys , us , pt s ≤ t − 1} ∪ {yt }. Here pt is an i.i.d. Bernoulli process such
that pt = 1 with probability p.
The initial state has a Gaussian distribution, with zero mean and variance E[x20 ], which we denote by ν0 . We wish to
find for some r > 0:
X3
inf J(x0 , γ) = Eνγ0 [ x2t + ru2t ],
γ
t=0

Compute the optimal control policy and the optimal cost. It suffices to provide a recursive form.
Hint: Show that the optimal control has a separation structure. Compute the conditional estimate through a revised
Kalman Filter due to the presence of pt .

Exercise 6.7.2 Let X, Y be Rn and Rm valued zero-mean random vectors defined on a common probability space,
which have finite covariance matrices. Suppose that their probability measures are given by PX and PY respectively.
Find
inf E[(X − KY )T (X − KY )],
K

that is find the best linear estimator of X given Y and the resulting estimation error.
Hint: You may pose the problem as a Projection Theorem problem.

Exercise 6.7.3 (Optimal Machine Repair) Consider a POMDP given by the following description. Let there be two
possible states that a machine can take: X = {0, 1}, where 0 is the bad (‘system is down’) state and 1 is the good
state. Let U = {0, 1}, where 0 is the ‘do nothing’ control and 1 is the ‘repair’ control. Suppose that the transition
probabilities are given by:

P (Xt+1 = 1|Xt = 1, Ut = 0) = 1 − η1 , P (Xt+1 = 0|Xt = 1, Ut = 0) = η1 > 0


P (Xt+1 = 1|Xt = 1, Ut = 1) = 1 − η2 , P (Xt+1 = 0|Xt = 1, Ut = 1) = η2 > 0
P (Xt+1 = 1|Xt = 0, Ut = 0) = 0, P (Xt+1 = 0|Xt = 0, Ut = 0) = 1
P (Xt+1 = 1|Xt = 0, Ut = 1) = α > 0, P (Xt+1 = 0|Xt = 0, Ut = 1) = 1 − α (6.46)

Thus, η1 is the failure probability when the state is good (and no repair) and η2 is the failure probability when the state
is good (and when there is repair) with η1 > η2 , and α is the success probability in the event of a repair.
The controller has access only to {0, 1}-valued measurement variables Y0 , · · · , Yt and U0 , · · · , Ut−1 , at time t, where
the measurements are generated by a binary symmetric channel:

(Y = x|X = x) = 1 − ϵ, P (Y = 1 − x|X = x) = ϵ,

for all x ∈ {0,.


The per-stage cost function c(x, u) is given by c(0, 0) = C, c(1, 0) = 0, c(0, 1) = c(1, 1) = R with 0 < R < C. Show
that there exists an optimal control policy for both finite-horizon as well as infinite horizon discounted cost problems.
140 6 Partially Observed Markov Decision Processes, Non-Linear Filtering, and the Kalman Filter

Exercise 6.7.4 (Zero-Delay Source Coding) Let {xt }t≥0 be an X-valued discrete-time Markov process where X can
be a finite set or Rn . Let there be an encoder which encodes (quantizes) the source samples and transmits the encoded
versions to a receiver over a discrete noiseless channel with input and output alphabet M = {1, 2, . . . , M }, where M
is a positive integer. The encoder policy γ is a sequence of functions {κt }t≥0 with κt : Mt × (X)t+1 → M. At time t,
the encoder transmits the M-valued message
qt = κt (It )
with I0 = x0 , It = (q[0,t−1] , x[0,t] ) for t ≥ 1, where. The collection of all such zero-delay encoders is called the
set of admissible quantization policies and is denoted by ΓA . A zero-delay receiver policy is a sequence of functions
γ d = {γtd }t≥0 of type γtd : Mt+1 → U, where U denotes the finite reconstruction alphabet. Thus

ut = γtd (q[0,t] ), t ≥ 0.

For the finite horizon setting the goal is to minimize the average cumulative cost (distortion)
 T −1 
d 1 X
Jπ0 (γ, γ d , T ) = Eπγ,γ c0 (xt , ut ) , (6.47)
0
T t=0

d
for some T ≥ 1, where c0 : X × U → R is a nonnegative cost (distortion) function, and Eπγ,γ 0
denotes expectation
with initial distribution π0 for x0 and under the quantization policy γ and receiver policy γ d .
a) Show that an optimal encoder uses a sufficient statistic, in particular, it uses P (dxt |q[0,t−1] ) and the time informa-
tion, for optimal performance.
b) Show that, when {xt } is i.i.d., any encoder and decoder pair can be replaced with one which only uses xt , that is:

qt = κt (xt )

and the decoder only uses


ut = γtd (qt ), t ≥ 0.

See [330], [321], [305] for finite sources and [343] for real sources and further relevant discussions, among many
other recent references.

Exercise 6.7.5 Let there be two decision makers, DM1 and DM2. Suppose that DMi (i = 1, 2) has access to:

Yi =X +Vi

where X, V 1 , V 2 are independent Gaussian random variables with unit variance and zero mean.
a) Find E[(X − E[X|Y i ])2 ] for i = 1, 2.
b) Suppose that DM1 and DM2 share their data Y 1 and Y 2 . Find

E[(X − E[X|Y 1 , Y 2 ])2 ]

c) Suppose that DM1 and DM2 share with each other their estimates E[X|Y i ]. That is, DM1 has access to Y 1 and
E[X|Y 2 ]; and DM2 has access to Y 2 and E[X|Y 1 ]. Find
  2    2 
1 2 2 1
E X − E X Y , E[X|Y ] and E X − E X Y , E[X|Y ]
7

The Average Cost Problem

In this chapter, we consider the following average cost problem of finding


T −1
∗ 1 γ X
J∞ (x) := inf J∞ (x, γ) = inf lim sup Ex [ c(xt , ut )] (7.1)
γ γ∈ΓA T →∞ T t=0

This is an important problem in applications where one is concerned about the long-term behaviour, unlike the dis-
counted cost setup where the primary interest is in the short-term time stages.
For the study of the average cost problem, we will follow three distinct approaches; the first two will be based on the
arrival at what we will call as the average cost optimality equation. The third approach will be based on the properties
of expected (or sample path) occupation measures and their limit behaviours, leading to a linear program involving
the space of probability measures. These approaches are related (e.g. via a dual optimization analysis [165, Chapter
12, p. 221], or a more direct stochastic analysis [160, Theorem 5.3]), however the conditions leading to solutions
under these approaches are not identical, therefore, the corresponding conditions of existence and structural results for
optimal policies are slightly different. As such, it will be instructive to study both approaches separately, as we do in
the following.

7.1 Average Cost and the Average Cost Optimality Equation (ACOE) or Inequality (ACOI)

To study the average cost problem, one approach is to establish the existence of an Average Cost Optimality Equation
(ACOE), and an associated verification theorem.

Definition 7.1.1 The collection of functions g : X → R, h : X → R, f : X → U is a canonical triplet if for all x ∈ X,


Z
g(x) = inf g(x′ )T (dx′ |x, u)
u∈U
 Z 
′ ′
g(x) + h(x) = inf c(x, u) + h(x )T (dx |x, u)
u∈U

with Z
g(x) = g(x′ )T (dx′ |x, f (x))
 Z 
′ ′
g(x) + h(x) = c(x, f (x)) + h(x )T (dx |x, f (x))

We will refer to these relations as the Average Cost Optimality Equation (ACOE).
142 7 The Average Cost Problem

Theorem 7.1.1 [Verification Theorem] Let g, h, f be a canonical triplet. a) If g is a constant and Exγ [|h(xn )|] < ∞
with
1 γ
lim sup E [h(xn )] = 0, (7.2)
n→∞ n x

for all x and under every policy γ, then the stationary deterministic policy γ ∗ = {f, f, f, · · · } is optimal so that

g = J∞ (x, γ ∗ ) = inf J∞ (x, γ)


γ∈ΓA

where
T −1
1 γ X
J(x, γ) = lim sup E [ c(xt , ut )].
T →∞ T x
k=0

Furthermore, if lim supn→∞ n1 |Exγ [h(xn )]| = 0,


n  
1 γ∗ X 1 γ∗
lim Ex [c(xt−1 , ut−1 )] − g ≤ lim sup |Ex [h(xn )] − h(x)| = 0 (7.3)
n→∞ n n→∞ n
t=1

b) If g, considered above, is not a constant and depends on x, then under any policy γ
N −1 T −1
1 γ∗ X 1 X
lim sup Ex [ g(xt )] ≤ inf lim sup Exγ [ c(xt , ut )],
N →∞ N t=0
γ N →∞ N
t=0

provided that (7.2) holds. Furthermore, γ ∗ = {f } is optimal.

Proof: We prove (a); (b) follows from a similar reasoning. For any admissible policy γ,
Z
E γ [h(xt )|x[0,t−1] , u[0,t−1] ] = h(y)P (xt ∈ dy|xt−1 , ut−1 ) (7.4)
y
Z
= c(xt−1 , ut−1 ) + h(y)P (dy|xt−1 , ut−1 ) − c(xt−1 , ut−1 ) (7.5)
y
 Z 
≥ min c(xt−1 , ut−1 ) + h(y)P (dy|xt−1 , ut−1 ) − c(xt−1 , ut−1 ) (7.6)
ut−1 ∈U y
= g + h(xt−1 ) − c(xt−1 , ut−1 ) (7.7)

Observe that with hM → h a monotonically increasing sequence of bounded functions hM which pointwise converges
to h, we have that

E γ [h(xt )] = lim E γ [hM (xt )]


M →∞
= lim E[E [h (xt )|x[0,t−1] , u[0,t−1] ]] = E[ lim E γ [hM (xt )|xt−1 , ut−1 ]]
γ M
M →∞ M →∞
≥ E[g + h(xt−1 ) − c(xt−1 , ut−1 )] (7.8)

Hence, for any admissible policy γ, x ∈ X, by re-arranging the terms, we have that for all n ∈ N
 n 
1 γ X
0 ≤ Ex h(xt ) − g − h(xt−1 ) + c(xt−1 , ut−1 )
n t=1

and thus
n
1 γ 1 γ 1 γX
g ≤ Ex [h(xn )] − Ex [h(x0 )] + Ex [ c(xt−1 , ut−1 )].
n n n t=1
7.1 Average Cost and the Average Cost Optimality Equation (ACOE) or Inequality (ACOI) 143

Taking the limit and using (7.2), we observe that g is a lower bound on the cost under any policy.
The above hold with equality if γ ∗ = {f } is adopted since γ ∗ provides the pointwise minimum. Thus, equality holds
under γ ∗ so that
∗ ∗ n
E γ [h(xn )] Exγ [h(x0 )] 1 γ∗ X
g= − + Ex [c(xt−1 , ut−1 )].
n n n t=1

Under (7.2),
n
1 γ∗ X
g = lim Ex [ c(xt−1 , ut−1 )].
n→∞ n
t=1

and
n  
1 γ∗ X 1 γ∗
E [ c(xt−1 , ut−1 )] − g ≤ |Ex [h(xn )]| + |h(x)| → 0,
n x t=1 n
as n → ∞ ⋄

Theorem 7.1.2 [Optimality Through Finite Horizon Limits] If γ ∗ = {f, f, f, · · · } is so that

J∞ (x, γ ∗ ) = lim sup inf J T (x, γ)


T →∞ γ∈ΓA

with
T −1
1 γ X
J T (x, γ) = Ex [ c(xt , ut )], (7.9)
T t=0

then γ ∗ is optimal.

Proof. The proof follows from the observation (by Lemma 5.5.1)

J∞ (x) ≥ lim sup inf J T (x, γ),
T →∞ γ∈ΓA

and that γ ∗ achieves this lower bound. ⋄

Remark 7.1. Recall that we had utilized the argument used in the proof of Theorem 7.1.2 while studying average cost
LQG problems in Theorem 5.12)(iii).

Remark 7.2. Note that if we have that, in Theorem 7.1.1,

g = lim sup inf J T (x, γ),


T →∞ γ∈ΓA

then it suffices to check (7.2) only for the policy γ ∗ to certify its optimality as this would ensure that the condition in
Theorem 7.1.2, on the achievability of g by g = J(x, γ ∗ ), is attained. Note the analogy with Lemma 5.5.4.

Definition 7.1.2 Let g be a constant and h : X → R, f : X → U be so that for all x ∈ X,


 Z 
′ ′
g + h(x) ≥ inf c(x, u) + h(x )T (dx |x, u) (7.10)
u∈U

Alternatively, let
 Z 
′ ′
g + h(x) ≤ inf c(x, u) + h(x )T (dx |x, u) (7.11)
u∈U

with, in either case


144 7 The Average Cost Problem
 Z   Z 
inf c(x, u) + h(x′ )T (dx′ |x, u) = c(x, f (x)) + h(x′ )T (dx′ |x, f (x))
u∈U

We will refer to (7.10) as the Average Cost Optimality Inequality (ACOI).

See, e.g., [14, Theorem 6.6] for the following:

Theorem 7.1.3 [Verification Theorem]


(i) Let (7.11) hold. If
1 γ
lim sup E [h(xn )] ≤ 0, (7.12)
n→∞ n x
for all x and under every policy γ. Then g is a lower bound under any policy.
(ii) On the other hand, with (7.10) so that
Z
g + h(x) ≥ c(x, f (x)) + h(x′ )T (dx′ |x, f (x))

and
1 γ∗
lim inf E [h(xn )] ≥ 0, (7.13)
n→∞ n x
holding with γ ∗ = {f, f, f, · · · }. Then the stationary deterministic policy γ ∗ = {f, f, f, · · · } satisfies

g ≥ J(x, γ ∗ )

Proof: For (i): for any policy γ,


Z
γ
E [h(xt )|x[0,t−1] , u[0,t−1] ] = h(y)P (xt ∈ dy|xt−1 , ut−1 ) (7.14)
y
Z
= c(xt−1 , ut−1 ) + h(y)P (dy|xt−1 , ut−1 ) − c(xt−1 , ut−1 ) (7.15)
y
 Z 
≥ min c(xt−1 , ut−1 ) + h(y)P (dy|xt−1 , ut−1 ) − c(xt−1 , ut−1 ) (7.16)
ut−1 ∈U y
≥ g + h(xt−1 ) − c(xt−1 , ut−1 ), (7.17)

where the last inequality is due to (7.11).


As in (7.8), we have that
E γ [h(xt )] ≥ E[g + h(xt−1 ) − c(xt−1 , ut−1 )]
Hence, for any policy γ
n
1 γX
0≤ E [h(xt ) − g − h(xt−1 ) + c(xt−1 , ut−1 )]
n x t=1
and
n
1 γ 1 1 X
g≤ Ex [h(xn )] − Exγ [h(x0 )] + Exγ [ c(xt−1 , ut−1 )].
n n n t=1

Taking the limit, we observe that g is a lower bound on the cost under any policy under (7.13).
For (ii): if we start the analysis above leading to (7.17) with γ ∗ , we have
7.2 The Value Iteration and Contraction Approach to the Average Cost Problem 145
n
X 

γ∗
Exγ h(xt ) − E [h(xt )|x[0,t−1] , u[0,t−1] ] = 0
t=1

and
Z

E γ [h(xt )|x[0,t−1] , u[0,t−1] ] = h(y)P (xt ∈ dy|xt−1 , ut−1 ) (7.18)
y
Z
= c(xt−1 , f (xt−1 )) + h(y)P (dy|xt−1 , f (xt−1 )) − c(xt−1 , f (xt−1 )) (7.19)
y
≤ g + h(xt−1 ) − c(xt−1 , f (xt−1 )) (7.20)

Iterating the above and dividing by n, we arrive at


n
1 γ∗ 1 ∗ 1 ∗X
g− Ex [h(xn )] + Exγ [h(x0 )] ≥ Exγ [ c(xt−1 , ut−1 )].
n n n t=1

Taking the limsup on both sides (and replacing lim sup with lim inf by reversing the negative sign on the left), and
(7.13) holding for γ ∗ = {f, f, f, · · · }, we establish the desired bound. ⋄

7.2 The Value Iteration and Contraction Approach to the Average Cost Problem

Fix z ∈ X and consider the space of measurable and bounded functions h with the restriction that h(z) = 0. Let
(g, h, f ) be a canonical triplet with g ≡ ρ ∈ R so that
 Z 
ρ + h(x) = inf c(x, u) + h(x′ )T (dx′ |x, u)
u∈U

7.2.1 Contraction under the span semi-norm

Consider the following assumption.

Assumption 7.2.1 For some α ∈ [0, 1), and for all x, x′ ∈ X and u, u′ ∈ U

∥P (·|x, u) − P (·|x′ , u′ )∥T V ≤ 2α

A sufficient condition for the above is the following minorization condition.

Assumption 7.2.2 There exists a positive measure µ′ with T (B|x, u) ≥ µ′ (B), for all B ∈ B(X) and all (x, u) ∈
X × U.

Observe that Assumption 7.2.1 is weaker than Assumption 7.2.2.


Consider the following span semi-norm:

∥u∥sp = sup u(x) − inf u(x)


x x

The space of measurable bounded functions that satisfy h(z) = 0 under the semi-norm ∥u∥sp is a Banach space (and
hence the semi-norm becomes a norm in this space since ∥u∥sp = 0 implies u ≡ 0).
Define
146 7 The Average Cost Problem
 Z 
T(h)(x) = inf c(x, u) + h(x′ )T (dx′ |x, u) (7.21)
u∈U

Let
(Tz (h))(x) = (T(h))(x) − (T(h))(z)
Note that Tz maps the aforementioned Banach space to itself under the measurable selection conditions reviewed in
Chapter 5. Under Assumption 7.2.2, and the measurable selection conditions reviewed in Chapter 5, we will show
(through similar steps as those in Chapter 5) that the map is a contraction:
First note that for pairs (x, u) and (x′ , u′ ), with µ(dx1 ) := P (dx1 |x, u) − P (dx1 |x′ , u′ ) defining a signed measure,
by the Jordan-Hahn decomposition theorem [180, Theorem 2.8] there exists A with µ(A) = −µ(Ac ) ≥ 0 so that the
restriction of µ to A (i.e., µA (B) := µ(B ∩ A) for every Borel B) defines a non-negative measure and the restriction
of −µ to Ac defines a non-negative measure with µ(A) − µ(Ac ) = ∥µ∥T V ≤ 2α and thus µ(A) ≤ α. Thus, for any
x, x′ , u, u;
Z
h(x1 )P (dx1 |x, u) − h(x1 )P (dx1 |x′ , u′ )
Z Z
= h(x1 )(P (dx1 |x, u) − P (dx1 |x′ , u′ )) + h(x1 )(P (dx1 |x, u) − P (dx1 |x′ , u′ ))
A Ac
Z Z
′ ′
= h(x1 )(P (dx1 |x, u) − P (dx1 |x , u )) − h(x1 )(P (dx1 |x′ , u′ ) − P (dx1 |x, u))
A Ac
Z   Z  
′ ′
≤ sup h(x1 ) (P (dx1 |x, u) − P (dx1 |x , u )) − inf h(x1 ) (P (dx1 |x′ , u′ ) − P (dx1 |x, u))
A x1 Ac x1
   
≤ sup h(x1 ) µ(A) − inf h(x1 ) ((−µ)(Ac ))
x1 x1

≤ α∥h∥sp

Then, note that for v1 and v2 bounded and with T(vi )(x) achieved with control uxi at x, we have that for any x, x′
   
′ ′
(T(v1 ))(x) − (T(v2 ))(x) − (T(v1 ))(x ) − (T(v2 ))(x )
Z  
x ′ x′
≤ (v1 (x1 ) − v2 (x1 )) P (dx1 |x, u2 ) − P (dx1 |x , u1 ) ≤ α∥v1 − v2 ∥sp ,

(where for the first term (T(v1 ))(x) − (T(v2 ))(x)we upper bound the difference by applying ux2 for the right hand side

of (7.21) involving v1 , and for the second term (T(v1 ))(x′ ) − (T(v2 ))(x′ ) we apply ux1 for bounding (T(v2 ))(x′ )) and
thus, since x, x′ are arbitrary, we have

∥T(v1 ) − T(v2 )∥sp ≤ α∥v1 − v2 ∥sp .

Furthermore, since (T(vi ))(z) only serves as a shift term in the following, we have that

∥Tz (v1 ) − Tz (v2 )∥sp = ∥T(v1 ) − T(v2 )∥sp ≤ α∥v1 − v2 ∥sp .

We can then state the following.

Theorem 7.2.1 [161, Lemma 3.5] The iterations

hn+1 = Tz (hn ),

with h0 ≡ 0 converges to a fixed point: Tz (h) = h, which leads to the ACOE triplet in Definition 7.1.1. In particular,
if the cost is bounded, under Assumption 7.2.1, and the controlled kernel satisfies the measurable selection conditions
given in Assumption 5.2.1 or 5.2.2, there exists a solution to the ACOE, which in turn leads to an optimal policy.
7.3 The Vanishing Discounted Cost Approach to the Average Cost Problem 147

7.2.2 Contraction under sup norm via minorization by equivalence with a discounted cost problem

We now present a more direct approach. Under the slightly stronger Assumption 7.2.2, we have that with

T ′ (·|x, u) = T (·|x, u) − µ′ (·)

a positive measure, the map


 Z 
(T′ (h))(x) = min c(x, u) + h(x1 )T ′ (dx1 |x, u)
u∈U

is a contraction (see [161, p.61] for a historical review on this approach). With this approach, one can avoid the use of
the span semi-norm approach. Accordingly, one can apply the standard value iteration algorithm using T′ , following
the proof of Theorem 5.5.2. The limit equation
 Z   Z  Z
h(x) = min c(x, u) + h(x1 )T ′ (dx1 |x, u) = min c(x, u) + h(x1 )T (dx1 |x, u) − h(x1 )µ′ (dx1 )
u∈U u∈U

h(x1 )µ′ (dx1 ).


R
is the desired ACOE in Definition 7.1.1 with g ≡

7.3 The Vanishing Discounted Cost Approach to the Average Cost Problem

7.3.1 Finite state and action spaces

Average cost emphasizes the asymptotic values of the cost function whereas the discounted cost emphasizes the short-
term cost functions. However, under technical restrictions, one can show that the limit as the discounted factor converges
to 1, one can obtain a solution for the average cost optimization. We now state one such condition below.

Theorem 7.3.1 [101] [47] [14, Theorem 4.3] Consider a controlled Markov chain where the state and action spaces
are finite, and suppose that under any stationary and deterministic policy the entire state space is a recurrent set. Let

X
Jβ (x) = inf Jβ (x, γ) = inf Exγ [ β t c(xt , ut )]
γ∈ΓA γ∈ΓA
t=0

and suppose that γn∗ is an optimal deterministic policy for Jβn (x). Then, there exists some γ ∗ ∈ ΓSD which is optimal
for every β sufficiently close to 1, and is also optimal for the average cost
T −1
1 γ X
J(x) = inf lim sup Ex [ c(xt , ut )]
γ∈ΓA T →∞ T t=0

P∞ policy f , by a slight change in notation from what was


Proof. First note that for every stationary and deterministic
considered earlier in the notes, Jβ (x, f ) := (1 − β)Exf [ k=0 β k c(x, f (xk ))] is a continuous function on [0, 1] (in β).
Let βn ↑ 1. For each βn , Jβn is achieved by a stationary and deterministic policy. Since there are only finitely many
such policies, there exists at least one policy f ∗ which is optimal for infinitely many βn ; call such a sequence βnk . We
will show that this policy is optimal for the average cost problem also.
It follows that (1 − βnk )Jβnk (x, f ∗ ) ≤ (1 − βnk )Jβnk (x, γ) for all γ. Then, infinitely often for every deterministic
stationary policy f :
(1 − βnk )Jβnk (x, f ∗ ) − (1 − βnk )Jβnk (x, f ) ≤ 0
We now claim that for some β ∗ < 1, Jβ (x, f ∗ ) ≤ Jβ (x, γ) for all β ∈ (β ∗ , 1). The function (1 − βnk )Jβnk (x, f ∗ ) −
(1 − βnk )Jβnk (x, f ) is continuous in β and uniformly bounded, therefore if the claim were not correct, the function
must have infinitely many zeros. On the other hand, one can write the equation
148 7 The Average Cost Problem
X
Jβ (x, f ) = c(x, f ) + β P (x′ |x, f (x))Jβ (x′ , f )
x′

in matrix form to obtain Jβ (·, f ) = (I − βP (· · · |·, f (·)))−1 c(·, f (·)). It follows that, (1 − z)(Jz (x, f ∗ ) − Jz (x, f ))
is a rational function (that is, ratio of two polynomials with finite order) on the open unit disk (in the complex region)
|z| < 1, such a function can only have finitely many zeros (unless it is identically zero): this follows by studying
the inverse matrix (I − zP )−1 which is analytic inside the unit disk and if it is non-zero on the boundary of the unit
disk at z = 1, it has to be bounded away from zero in a neighborhood of z = 1 inside the unit disk. Therefore, it
must be that for some β ∗ < 1, Jβ (x, f ∗ ) ≤ Jβ (x, γ) for all β ∈ (β ∗ , 1). We note here that such a policy is called a
Blackwell-Optimal Policy. Now,

(1 − βnk )Jβnk (x, f ∗ ) ≤ (1 − βnk )Jβnk (x, γ) (7.22)

for any γ and thus,


T −1
1 f∗ X
J(x, f ∗ ) = lim inf E [ c(xk , uk )] ≤ lim inf (1 − βnk )Jβnk (x, f ∗ ) = lim sup(1 − βnk )Jβnk (x, f ∗ )
T →∞ T x nk →∞ nk →∞
k=0
T −1
1 X
≤ lim sup(1 − βnk )Jβnk (x, γ) ≤ lim sup Exγ [ c(xk , uk )] (7.23)
nk →∞ T →∞ T
k=0

In the first equality, we use the fact that the limit exists. In the above, the sequence of inequalities follow from the
following Abelian inequalities (see [164, Lemma 5.3.1]): Let an be a sequence of non-negative numbers and β ∈ (0, 1).
Then,
N −1 ∞
1 X X
lim inf am ≤ lim inf (1 − β) β m am
N →∞ N m=0 β↑1
m=0
∞ N −1
X 1 X
≤ lim sup(1 − β) β m am ≤ lim sup am (7.24)
β↑1 m=0 N →∞ N m=0

As a result, f ∗ is optimal. The optimal cost does not depend on the initial state by the recurrence condition and
irreducibility of the chain under the optimal policy. ⋄
In the following, we consider more general state spaces and generalize the results presented above.

7.3.2 Standard Borel state and action spaces, ACOE and ACOI

Consider the value function for a discounted cost problem as discussed in Section 5.5:
 Z 
Jβ (x) = min c(x, u) + β Jβ (y)T (dy|x, u) , x ∈ X. (7.25)
u∈U X

Let x0 be an arbitrary state and for all x ∈ X consider

Jβ (x) − Jβ (x0 )
 Z 
′ ′
= min c(x, u) + β T (dx |x, u)(Jβ (x ) − Jβ (x0 )) − (1 − β)Jβ (x0 )
u∈U

As discussed in Section 5.5, this has a solution for every β ∈ (0, 1) under measurable selection conditions.
Arriving at the Average Cost Optimality Equation. We recall that a family of functions F mapping a metric space S
to R is said to be equicontinuous at a point x0 ∈ S if, for every ϵ > 0, there exists a δ > 0 such that d(x, x0 ) ≤ δ =⇒
|f (x) − f (x0 )| ≤ ϵ for all f ∈ F . The family F is said to be equicontinuous if it is equicontinuous at each x ∈ S.
7.3 The Vanishing Discounted Cost Approach to the Average Cost Problem 149

Now suppose that hβ (x) := Jβ (x) − Jβ (x0 ) is equicontinuous (over β) and X is compact. By the Arzela-Ascoli
Theorem (Theorem 7.3.2), taking β ↑ 1 along some sequence, for some subsequence, Jβnk (x) − Jβnk (x0 ) → η(x)
uniformly. If the cost is bounded, then, along a further subsequence,

(1 − βnk )Jβnk (x0 ) → ζ ∗ (7.26)

for some ζ ∗ (which is to be shown to be independent of x0 ). If we could also exchange the order of the minimum and
the limit, one obtains the Average Cost Optimality Equation (ACOE):
 Z 
′ ′ ∗
η(x) = min c(x, u) + T (dx |xt , ut )η(x ) − ζ , (7.27)
u∈U

which has the form of the equations in Definition 7.1.1.


We now make this observation formal (and relax the compactness assumption on the state space).

Assumption 7.3.1
(a) The one stage cost function c is bounded and continuous.
(b) The stochastic kernel T ( · |x, u) is weakly continuous in (x, u) ∈ X × U, i.e., if (xk , uk ) → (x, u), then
T ( · |xk , uk ) → T ( · |x, u) weakly.
(c) U is compact.
(d) X is σ-compact, that is, X = ∪n Sn where Sn ⊂ Sn+1 and each Sn is compact.

In addition to Assumption 7.3.1, we impose the following assumption in this section.

Assumption 7.3.2
There exists α ∈ (0, 1) and N ≥ 0, and a state z0 ∈ X such that,
(e) −N ≤ hβ (z) ≤ N for all z ∈ X and β ∈ [α, 1), where

hβ (z) = Jβ (z) − Jβ (z0 ),

for some fixed z0 ∈ X.


(f) The sequence {hβ(k) } is equicontinuous, where {β(k)} is a sequence of discount factors converging to 1 which

satisfies limk→∞ (1 − β(k))Jβ(k) (z) = ρ∗ for all z ∈ X for some ρ∗ ∈ [0, L].

Note that when the one stage cost function c is bounded by some L ∈ R+ , we must have

|(1 − β)Jβ∗ (z)| ≤ L

for all β ∈ (0, 1) and z ∈ X. Let us recall the Arzela-Ascoli theorem.

Theorem 7.3.2 [110] Let F be an equicontinuous family of functions on a compact space X and let hn be a sequence
in F such that the range of fn is compact. Then, there exists a subsequence hnk which converges uniformly to a
continuous function. If X is σ-compact, that is X = ∪n Kn with Kn ⊂ Kn+1 with Kn compact, the same result holds
where hnk converges pointwise to a continuous function, and and the convergence is uniform on compact subsets of X.

Theorem 7.3.3 Under Assumptions 7.3.1 and 7.3.2, there exist a constant ρ∗ ≥ 0, a continuous and bounded h from X
to R with −N ≤ h( · ) ≤ N , and {f ∗ } ∈ ΓS such that (ρ∗ , h, f ∗ ) satisfies the ACOE; that is,
Z

ρ + h(z) = min(c(z, u) + h(y)T (dy|z, u))
u∈U
ZX
= c(z, f ∗ (z)) + h(y)T (dy|z, f ∗ (z)),
X
150 7 The Average Cost Problem

for all z ∈ X. Moreover, {f ∗ } is optimal and ρ∗ is the value function, i.e.,

inf J(φ, z) =: J ∗ (z) = J({f ∗ }, z) = ρ∗ ,


φ

for all z ∈ X.

Proof. By (7.26), we have that (1 − βnk )Jβ (x0 ) → ρ∗ for some subsequence nk as βnk ↑ 1 and some ρ∗ . Observe that
for any x ∈ X
(1 − βnk )Jβnk (x) = (1 − βnk )(Jβnk (x) − Jβnk (x0 )) + (1 − βnk )Jβnk (x0 ),
which, by the uniform boundedness of Jβnk (x) − Jβnk (x0 ), implies that the limit ρ∗ does not depend on x. By As-
sumption 7.3.2-(f) and Theorem 7.3.2, there exists a further subsequence of hnk , {hβ(kl ) }, which converges (uniformly
on compact sets) to a continuous and bounded function h. Take the limit in (7.36) along this subsequence, i.e., consider
Z
ρ∗ + h(z) = lim min[c(z, u) + β(kl ) hβ(kl ) (y)T (dy|z, u)]
l U
ZX
= min lim[c(z, u) + β(kl ) hβ(kl ) (y)T (dy|z, u)]
U l X
Z
= min[c(z, u) + h(y)T (dy|z, u)].
U X

Here, somewhat similar to Lemma 5.2.2, the Rexchange of limit and minimum follows from writing (using the compact-
ness of U, the continuity of [c(z, u) + β(kl ) X hβ(kl ) (y)T (dy|z, u)] on U, and the equicontinuity of {hβ(k) }):
Z Z
min[c(z, u) + β(kl ) hβ(kl ) (y)T (dy|z, u)] = c(z, ul ) + β(kl ) hβ(kl ) (y)T (dy|z, ul )
U X X
Z Z
min[c(z, u) + h(y)T (dy|z, u)] = c(z, u∗ ) + h(y)T (dy|z, u∗ )
U X X
and showing that
 Z Z 
max (β(kl )hβ(kl ) (y) − h(y))T (dy|z, ul ) , (β(kl )hβ(kl ) (y) − h(y))T (dy|z, u∗ ) → 0. (7.28)
X X

The last item follows from a contrapositive argument. Suppose that the term does not converge to zero, implying
that for some subsequence it remains above some ϵ > 0 As in the proof of Lemma 5.2.2, for any such subsequence
we would have the following. By compactness of the action space U, we would have a further subsequence so that
un → u (ignoring the subscripts) for some u along this further subsequence. By weak continuity of the kernel Under
Assumption 7.3.1, we then have that T (dy|z, un ) → T (dy|z, u). Since hβ(kl ) converges uniformly on compact sets,
convergence to zero follows from Theorem D.3.1(i), concluding the argument (a more direct argument would be as
follows: Since for every {un → u}, the set of probability measures T (dy|z, u R n ) is tight, for every ϵ > 0 (by weak
continuity in Assumption 7.3.1), one can find a compact set Kn ⊂ X so that X\Kn hβ(kl ) (y)T (dy|z, u)] ≤ ϵ (here,
by Assumption 7.3.2-(e), uniform boundedness of hβ(kl ) is critical). Since on Kn , hβ(kl ) → h uniformly and h is
bounded, the result follows). ⋄

Remark 7.3. One can also consider (the slightly stronger condition of) Assumptions 4.2.1 and 5.5.1 of [164]; see e.g.
[164, Theorem 5.5.4])); see also [285, Theorem 3.8]. Further conditions also appear in the literature; see Hernandez-
Lerma and Lasserre [165] for a detailed analysis for the unbounded cost setup, and [90] for such results and a detailed
literature review. Further conditions are available in [146] [317], among other references.

Corollary 7.3.1 (Beyond Minorization) [99, Lemma 2.4] Consider Assumption 5.5.1 with K2 < 1 and that U and X
are compact. Then, Theorem 7.3.3 is applicable.

Proof. The proof follows since the equicontinuity condition is satisfied in view of (5.38) (see [270, Theorem 4.37]) for
all β ∈ (0, 1]. ⋄
7.3 The Vanishing Discounted Cost Approach to the Average Cost Problem 151

For an explicit proof, see [99, Lemma 2.4 and Lemma 2.5]. The above is particularly useful for belief-MDPs; see
Theorem 6.3.8. If compactness does not hold in Theorem 7.3.1 but X is σ-compact, then by Lipschitz regularity for
each compact restriction, uniform boundedness would apply and convergence of hβ along a subsequence to a limit,
uniformly on compact sets, would apply. To account for the non-boundedness of the limit h, an integrability condition
leading to (7.28) would be sufficient to arrive at the optimality equation.
In the following, we obtain two partial generalizations of Theorem 7.3.1 to the standard Borel space setup: First, we
observe that under the conditions of Theorem 7.3.3, since a solution to ACOE exists; every subsequential limit in (7.26)
will need to be identical. This leads to the following.

Theorem 7.3.4 Under the conditions of Theorem 7.3.3, (7.26) can be refined to

lim(1 − β)Jβ (x0 ) → ζ ∗ , (7.29)


β↑1

where Jβ is defined in (7.25) and ζ ∗ is the optimal average cost. That is, we will have sequential convergence (and not
just subsequential convergence).

The argument is via contraposition; suppose the limit of two subsequences were different. Each subsequence would
then have further subsequences which would satisfy the ACOE and via the verification theorem would lead to the
(same) optimal cost ζ ∗ . Hence the limits must be identical.
As a note, there is no claim that the limit of the relative value functions, h, is unique. If minorization holds (Assumption
7.2.2) then indeed h is unique up to a constant; in the absence of minorization, we do not have such a claim; see
also [353].
We showed in Theorem 7.3.4 that the value functions of discounted cost criteria converges to the value under the average
cost criterion under the conditions of Theorem 7.3.3. In practice, we would like to also see whether the policies solving
the discounted cost problem (or that are at least near optimal for the discounted cost problem) are (also) near optimal for
the average cost criterion. Such a result has significant implications for numerical and reinforcement learning theoretic
methods (e.g, [91, Theorem 5]).

Theorem 7.3.5 Let Assumptions 7.3.1 and 7.3.2 hold. Let γβ solve the discounted cost optimality equation (7.25).
Then, for every ϵ > 0, there exists β large enough such that

J(x, γβ ) − ζ ∗ < ϵ,

where ζ ∗ is the optimal average cost. That is, the discounted cost optimal policy is near-optimal for the average cost
criterion.

Proof. Write

hβ (x)
 Z 
= min c(x, u) + β T (dx′ |x, u)(hβ (x′ )) − (1 − β)Jβ (x0 )
u∈U
 Z 
= c(x, γβ (x)) + β T (dx′ |x, γβ (x))(hβ (x′ )) − (1 − β)Jβ (x0 ) (7.30)

It follows that
Z
hβ (x) + ζ ∗ − (ζ ∗ − (1 − β)Jβ (x0 )) + (1 − β) T (dx′ |x, γβ (x))(hβ (x′ ))
 Z 
′ ′
= c(x, γβ (x)) + T (dx |x, γβ (x))(hβ (x )) (7.31)
152 7 The Average Cost Problem

We have by assumption that hRβ is uniformly bounded. Now, take β sufficiently close to 1 so that |ζ ∗ −(1−β)Jβ (x0 )| ≤
ϵ by (7.29), and that |(1 − β) T (dx′ |x, γβ (x))(hβ (x′ ))| ≤ ϵ. Then, we have

hβ (x) + ζ ∗ + 2ϵ
 Z 
′ ′
≥ c(x, γβ (x)) + T (dx |x, γβ (x))(hβ (x )) (7.32)

Via Theorem 7.1.3(ii), since hβ is bounded, the above implies that γβ achieves an average cost not larger than ζ ∗ + 2ϵ.

We now establish near optimality of near optimal discounted solutions for the average cost setup.

Theorem 7.3.6 Let Assumptions 7.3.1 and 7.3.2 hold. Let βϵ be taken as in Theorem 7.3.5, be such that
ϵ
|ζ ∗ − (1 − βϵ )Jβϵ (x0 )| ≤
2
and
ϵ
(1 − βϵ )∥hβϵ ∥∞ ≤ ,
2
so that γβϵ is ϵ-optimal. Suppose that γβδ ϵ is such that

Jβϵ (x, γβδ ϵ ) − Jβϵ (x) < δ.

Then,
J(x, γβδ ϵ ) − ζ ∗ < ϵ + δ
That is, the near optimal discounted cost optimal policy is near-optimal for the average cost criterion as well.

Proof. We follow the proof of Theorem 7.3.5 with a minor variation as follows: Let us write the cost attained by the
policy γβδ ϵ as:
Jβϵ (x, γβδ ϵ ) = c(x, γβδ ϵ ) + βϵ E[Jβϵ (x1 , γβδ ϵ )|x0 = x, u0 = γβδ ϵ (x)]
This follows from the fact that the cost is bounded and the arguments used in Chapter 5 on the verification theorem;
alternatively see Section 8.1.2.
With x0 as in the proof of Theorem 7.3.5, write

h̄βϵ (x) := Jβϵ (x, γβδ ϵ ) − Jβϵ (x0 )

Then, we have that, by substitution,

h̄βϵ (x) = c(x, γβϵδ ) + βϵ E[h̄βϵ (x1 )|x0 = x, u0 = γβδ ϵ (x)] − (1 − βϵ )Jβϵ (x0 )

Then,
Z
h̄βϵ (x) + (1 − βϵ )Jβϵ (x0 ) + (1 − βϵ ) T (dx′ |x, γβδ ϵ (x))(h̄βϵ (x1 )) = c(x, γβδ ϵ ) + E[h̄βϵ (x1 )|x0 = x, u0 = γβδ ϵ (x)]

We have that |(1 − βϵ )Jβϵ (x0 ) − ζ ∗ | ≤ ϵ


2 and that ∥h̄βϵ ∥∞ ≤ supx∈X |Jβϵ (x, γβδ ϵ ) − Jβϵ (x)| + ∥hβϵ ∥∞ ≤ δ + ∥hβϵ ∥∞ .
This implies that

ϵ ϵ
h̄βϵ (x) + + (1 − βϵ )δ + ≥ c(x, γβϵδ ) + E[h̄βϵ (x1 )|x0 = x, u0 = γβδ ϵ (x)]
2 2

This has the same form as (7.32) and thus Theorem 7.1.3(ii) implies that γβδ ϵ achieves an average cost of no larger than
ζ∗ + ϵ + δ
7.3 The Vanishing Discounted Cost Approach to the Average Cost Problem 153


Average Cost Optimality Inequality. If one cannot verify the equicontinuity assumption or the boundedness condi-
tions, the following holds; note that the condition of strong continuity in actions for every fixed state is required here.
The result essentially follows from [164, Theorem 5.4.3] with some variations in the conditions.
R
Theorem 7.3.7 Let for every measurable and bounded g, the integral g(xt+1 )P (dxt+1 |xt = x, ut = u) be continu-
ous in u for every x, and there exist N < ∞ and a function b(x) with

−N ≤ hβ (x) ≤ b(x), β ∈ (0, 1), x ∈ X (7.33)

and for all β ∈ [α, 1) for some α < 1 and M ∈ R+ :

(1 − β)Jβ∗ (z) ≤ M. (7.34)

Under these conditions, the Average Cost Optimality Inequality (ACOI) holds for appropriate η, ζ ∗ , f :
 Z 
′ ′ ∗
η(x) ≥ min c(x, u) + T (dx |xt , ut )η(x ) − ζ
u∈U
 Z 
′ ′ ∗
= c(x, f (x)) + T (dx |x, f (x))η(x ) − ζ (7.35)

In particular, the stationary and deterministic policy γ = {f, f, f, · · · } is optimal.

Proof. By (7.34) and (7.26), we have that (1 − βnk )Jβnk (x0 ) → ζ ∗ for some subsequence nk and βnk ↑ 1. On the
other hand by the optimality of Jβ , and by Abelian inequalities, under any policy γ, and any sequence β ↑ 1
∞ T −1
X 1 γ X
lim sup(1 − β)Jβ (x0 ) ≤ lim sup(1 − β)Exγ0 [ β k c(xk , uk )] ≤ lim sup E [ c(xk , uk )]
β→1 β→1 T →∞ T x0
k=0 k=0

thus ζ ∗ is a lower bound under any admissible policy. We now make the argument that this applies for any initial
condition: Consider some arbitrary state x ∈ X,

(1 − βnk )Jβnk (x) = (1 − βnk )(Jβnk (x) − Jβnk (x0 )) + (1 − βnk )Jβnk (x0 )

By (7.33), (1 − βnk )(Jβnk (x) − Jβnk (x0 )) → 0. Thus, for any x, (1 − βnk )Jβnk (x) → ζ ∗ .
We now show that (7.35) holds. For this, consider again:

Jβ (x) − Jβ (x0 )
 Z 
= min c(x, u) + β T (dx′ |x, u)(Jβ (x′ ) − Jβ (x0 )) − (1 − β)Jβ (x0 ) (7.36)
u∈U

Observe the following along the subsequence nk , with hβ (x) = Jβ (x) − Jβ (x0 ), (1 − βnk )Jβnk (x0 ) → ζ ∗ and

lim inf Jβnk (x) − Jβnk (x0 )


nk →∞
 Z 
= lim inf min c(x, u) + βnk T (dx′ |x, u)hβnk (x′ ) − (1 − βnk )Jβnk (x0 )
nk →∞ u∈U
 Z 
= lim inf min c(x, u) + βnm T (dx′ |x, u)hβnm (x′ ) − (1 − βnk )Jβnk (x0 )
nk →∞ nm >nk u∈U
 Z 
≥ lim min c(x, u) + βnk T (dx′ |x, u)Hnk (x′ ) − (1 − βnk )Jβnk (x0 )
nk →∞ u∈U
 Z 
= min c(x, u) + T (dx′ |x, u)η(x′ ) − ζ ∗
u∈U
154 7 The Average Cost Problem

where Hnk (x) := min(inf nm >nk hβnm (x), nk ) (so that this is a bounded function) and η(x) = limm→∞ Hm (x) (so
that η(x) = lim inf nk →∞ Jβnk (x) − Jβnk (x0 ) which is the left hand term of the equation above). The last equality
holds since Hk ↑ η as shown inR Lemma 5.5.2 (though with an inequality sign since η is not necessarily bounded), and
bounded from below and that T (dx′ |x, u)Hk (x′ ) is continuous on U (this is where we use the strong continuity in
actions property given as a hypothesis).
We now show that the stationary policy γ = f ∞ =: {f, f, f, · · · } is optimal, via Theorem 7.1.3(ii): Using (7.35)
repeatedly, we have that
T −1
∞ X ∞
Exf [ c(xk , uk )] ≤ T ζ ∗ + η(x) − Exf [η(xT )] ≤ T ζ ∗ + η(x) + N.
k=0

Dividing by T and taking the lim sup, leads to the result that
T −1
1 f∞ X
lim sup E [ c(xk , uk )] ≤ ζ ∗ .
T →∞ T x
k=0

This completes the proof. ⋄


Further sufficient conditions exist in the literature for ACOE or ACOI to hold (see [165], [318]). These conditions
typically have the form of Assumption 5.5.2 or 5.5.3 together with geometric ergodicity conditions with condition
(5.40) replaced with conditions of the form:
Z
sup w(y)T (dy|x, u) ≤ αw(x) + Kϕ(x, u),
u∈U X

where α ∈ (0, 1), K < ∞ and ϕ a positive function. In some approaches, ϕ and w needs to be continuous, in others it
does not. For example if ϕ(x, u) = 1{x∈C} for some small set C, then we recover a condition similar to (4.29) leading
to geometric ergodicity.
We also note that for the above arguments to hold, there does not need to be a single invariant distribution. Here in
(7.36), the pair x and x0 should be picked as a function of the reachable set under a given sequence of policies. The
analysis for such a condition is tedious in general since for every β a different optimal policy will typically be adopted;
however, for certain applications the reachable set from a given point may be independent of the control policy applied.

7.4 The Convex Analytic Approach to Average Cost Markov Decision Problems

The convex analytic approach (typically attributed to Manne [221] and Borkar [59] (see also [164])) is a powerful ap-
proach to the optimization of infinite-horizon problems. It is particularly effective in proving results on the optimality
of stationary policies, which can lead to a linear program. This approach is particularly effective for constrained opti-
mization problems and infinite horizon average cost optimization problems. It avoids the use of dynamic programming
or iterative contraction methods.
We are interested in the minimization
T
1 γ X
inf lim sup Ex0 [ c(xt , ut )], (7.37)
γ∈ΓA T →∞ T t=1

where, as before, Exγ0 [·] denotes the expectation over all sample paths with initial state given by x0 under some admis-
sible policy γ.
7.4 The Convex Analytic Approach to Average Cost Markov Decision Problems 155

7.4.1 Finite state/action setup

We first consider the finite space setting where both X and U are finite sets. We study the limit distribution of the
following empirical occupation measures (and their expected values), under any policy γ in ΓA . Let for T ≥ 1
T −1
1 X
vT (D) = 1{(xt ,ut )∈D} , D ∈ B(X × U).
T t=0

Consider policy γ in ΓA , x0 ∼ η, and let for T ≥ 1, the expected empirical occupation measures be given with
−1
 TX 
1
µT (D) = Eηγ [vT (D)] = Eηγ 1{xt ,ut )∈D} , D ∈ B(X × U)
T t=0

Let for η ∈ P(X × U), X


ηT (A × U) := T (A|x, u)η(x, u).
x∈X,u∈U

We then have

(µT T )(A × U)
X
= T (A|x, u)µT (x, u)
x∈X,u∈U
X
= E γ [1{x1 ∈A} |x0 = x, u0 = u]µT (x, u)
x∈X,u∈U
T −1  
1 X γ X
= E E[1{xk+1 ∈A} |xk = x, uk = u]1{xk =x,uk =u}
T
k=0 x∈X,u∈U
T −1
1 X X
= E γ [1{xk+1 ∈A} |xk = x, uk = u]P (xk = x, uk = u)
T
k=0 x∈X,u∈U
T −1
1 X
= E γ [1{xk+1 ∈A} ] (7.38)
T
k=0

Then, through what is often referred to as a Krylov-Bogoliubov-type argument, for every A ⊂ X,

|µT (A × U) − µT T (A × U)|
−1
 TX T −1 
γ 1
X
= Eµ0 1{xt ,ut )∈(A×U)} − 1{xt+1 ,ut+1 )∈(A×U)} (7.39)
T t=0 t=0
1
≤ → 0,
T
as T → ∞. Notice that the above applies for any policy γ ∈ ΓA .
Now, if we can ensure that for some subsequence, µtk → µ for some probability measure µ, it would follow that
µtk T (A × U) → µT (A × U).
Now define
 X 
G= v ∈ P(X × U) : v(B × U) = P (xt+1 ∈ B|xt = x, ut = u)v(x, u), B ∈ B(X)
x,u
156 7 The Average Cost Problem

Further define
 X 
GX = v ∈ P(X × U) : ∃γ ∈ ΓS , v(A) = P γ ((xt+1 , ut+1 ) ∈ A|xt = x, ut = u)v(x, u), A ∈ B(X × U)
x,u
(7.40)

We can establish the equivalence of these sets of measures: It is evident that GX ⊂ G since there are (seemingly) fewer
restrictions for G. We can show that these two sets are indeed equal: P
For v ∈ G, if we write: v(x, u) = π(x)η(u|x) for
some η, then, we can construct a consistent v ∈ GX : v(B × C) = x∈B η(C|x)π(x). The set G is called the set of
invariant occupation measures (or, as is used more commonly in the literature: ergodic occupation measures).
Thus, every converging subsequence µtk will converge to G. And hence, any sequence {µk } will have a converging
subsequence whose limit will be in the set G. This is where finiteness is helpful: If the state space were countable,
there would be no guarantee that every sequence of occupation measures would have a converging subsequence. The
following has thus been established.

Lemma 7.4.1 Under any admissible policy, any converging subsequence {µtk } will converge to the set G.
P
Let ⟨µ, c⟩ := µ(x, u)c(x, u). Let us again write that

J ∗ (x) := inf J(x, γ),


γ∈ΓA

with
T −1
1 γ X
J(x, γ) := lim sup Ex [ c(xt , ut )],
T →∞ T t=0
or
J(x, γ) := lim sup⟨µT , c⟩,
T →∞

where µT is the expected empirical occupation measure under γ. Now, we have that, for any policy γ,

lim sup⟨µT , c⟩ ≥ lim inf ⟨µT , c⟩


T →∞ T →∞

= lim ⟨µTk , c⟩ = ⟨ lim



µTk′ , c⟩ ≥ δ ∗ (7.41)
Tk →∞ Tk →∞

where X
δ ∗ = inf v(x, u)c(c, u)
v∈G

In the above, µTk is a subsequence for which ⟨µTk , c⟩ converges to the liminf value, there is a further subsequence µTk′
(due to the compactness of the space of expected empirical occupation measures) which has a limit, and this limit is in
G.
That is,
lim ⟨µTk , c⟩ = ⟨ lim

µTk′ , c⟩ ≥ δ ∗ .
Tk →∞ Tk →∞

Thus, we have established that


J ∗ (x) ≥ δ ∗
If the initial state, or measure on the initial state, can be selected appropriately, or if the controlled Markov chain under
an optimal policy is positive Harris recurrent the above also becomes an equality. The solution to this problem then
gives us the optimal cost (under any policy). Thus, an optimal policy can be obtained through the following linear
program:
Linear Program For Finite Models.
7.4 The Convex Analytic Approach to Average Cost Markov Decision Problems 157

Given a cost function c and transition kernel T , find the minimum


X
min ν(x, u)c(x, u). (7.42)
ν∈G
X×U

over all probability measures ν that satisfy


 X 
ν ∈ G = µ ∈ P(X × U) : µ(z, U) = T (z|(x, u))µ(x, u), z∈X .
X×U

where the constraint set can also be written as


X X
µ(z, j) = T (z|(x, u))µ(x, u), z∈X
j X×U

with
µ(x, u) ≥ 0, x ∈ X, u ∈ U
X
µ(x, u) = 1
x∈X,u∈U

All of these are linear/affine constraints.


If under any stationary policy the induced Markov chain would be irreducible, then the solution of the problem (7.42)

above equals the optimal cost J∞ (x). Let µ∗ be the optimal occupation measure (this exists since the state space is
u)c(x, u) is continuous in µ). This induces an optimal policy γ ∗ (u|x) as
P
finite, and thus G is compact, and X×U P µ(x,∗
(defined almost surely, i.e., for x with U µ (x, u) > 0):

µ∗ (x, u)
γ ∗ (u|x) = P ∗ .
U µ (x, u)

Thus, we can find the optimal policy through a linear program.

7.4.2 General state/action spaces under weak continuity

The arguments presented above apply to general spaces as well. However, for the more general case considered here,
we need to ensure that the set of expected occupation measures is tight, and that the set G is closed. In the following,
we follow the presentation in [16].
We first study the weakly continuous setup, studied in [14, 14, 59, 160, 164, 200], on the existence of an optimal µ ∈ G
under the hypothesis that the transition kernel T is weakly continuous (Assumption 5.2.1(i)).

(H1)The transition kernel T is weakly continuous, that is


Z
T (f )(x) := f (z)T (dz|x, u)
X

is continuous in x, u for all f ∈ Cb (X). Recall that this is the same as Assumption 5.2.1(i).

Continuing, for T ≥ 1, we let


T −1
1 X
vT (D) = 1D (Xt , Ut ), D ∈ B(X × U) .
T t=0

Consider any policy γ in ΓA , X0 ∼ ν, and let for T ≥ 1,


158 7 The Average Cost Problem
"T −1 #
1 γ X
µγT (D) = Eνγ [vT (D)] = Eν 1D (Xt , Ut ) , D ∈ B(X × U) .
T t=0

We refer to µγT T >0 as the family of mean empirical occupation measures under the policy γ ∈ ΓA , and with initial


distribution ν. Again, through a Krylov-Bogoliubov-type argument, for every A ∈ B(X), we have


"T −1 T
#
γ γ 1 γ X X
|µT (A × U) − µT T (A)| = |E 1A×U (Xt , Ut ) − 1A×U (Xt , Ut ) |
T ν t=0 t=1 (7.43)
1
≤ → 0 as T → ∞ .
T

 for any policy γ ∈ ΓA . Suppose that, along some


Observe that (7.43) holds subsequence {tk } ⊂ N, µγt converges
γ
weakly to some µ ∈ ( (x, u) : x ∈ X, u ∈ U(x) ), which we denote as µtk ⇒ µ. We write the triangle inequality

|µ(f ) − µT (f )| ≤ |µ(f ) − µγtk (f )| + |µγtk (f ) − µγtk T (f )| + |µγtk T (f ) − µT (f )| (7.44)



for f ∈ Cb (X). This notation is consistent since f may be viewed also as an element of Cb ( (x, u) : x ∈ X, u ∈
U(x) ). Suppose that Assumption 5.2.1(i)) holds. The first term on the right hand side of (7.44) vanishes as k → ∞
by weak convergence, while the second term does the same by (7.43). Since

µγtk T (f ) = µγtk (T f ), (7.45)



and T f ∈ Cb ( (x, u) : x ∈ X, u ∈ U(x) ) by Assumption 5.2.1(i)), it follows that the third term also vanishes
as k → ∞ by the weak convergence µγtk ⇒ µ. Since the class Cb (X) distinguishes points in P(X), this shows that
µ(A, U) = µT (A) for all A ∈ B(X), which implies that µ ∈ G by the definition of the latter. Thus we have shown the
following.

Lemma 7.4.2 Under Assumption 5.2.1(i)), the limit of any weakly converging subsequence of mean empirical occupa-
tion measures is in G.

This expected average cost can be written as

J(x, γ) = lim sup ⟨µγT , c⟩ ,


T →∞

where µγT is the mean empirical occupation measure under γ. Let {tk } ⊂ N be a subsequence along which ⟨µγtk , c⟩
converges to J(x, γ) and suppose that µtk ⇒ µ ∈ G. Then
D E
J(x, γ) = lim inf ⟨µγtk , c⟩ ≥ lim µγtk , c = ⟨µ, c⟩ ≥ δ ∗ , (7.46)
tk →∞ tk →∞

where for the first inequality we use the fact that, since c is lower semi-continuous (l.s.c.) and bounded from below, the
map µ → ⟨µ, c⟩ is lower semi-continuous. The above shows that J ∗ (x) ≥ δ ∗ . We now establish conditions for which
the above is indeed an equality.

Assumption 7.4.1 (A) The state and action spaces X and U are Polish. The set-valued map U : X → B(U) is upper
semi-continuous and closed-valued.
(A’)The state and action spaces X and U are compact. The set-valued map U : X → B(U) is upper semi-continuous
and closed-valued.

 running cost function c(x, u) is is l.s.c. and c : (x, u) : x ∈ X, u ∈ U(x) → R is inf-compact,
(B) The non-negative
i.e. {(x, u) ∈ (x, u) : x ∈ X, u ∈ U(x) : c(x, u) ≤ α} is compact for every α ∈ R+ .
(B’)The cost function c is bounded and l.s.c..
7.4 The Convex Analytic Approach to Average Cost Markov Decision Problems 159

(C) There exists a policy and an initial state leading to a finite cost η ∈ R+ .
(D) Assumption 5.2.1(i)) holds.
(E) Under every stationary policy, the induced Markov chain is Harris recurrent.

Before we present a theorem, we recall the discussion


R in Section 3.4.1 concerning ergodic properties of (control-free)
Markov chains: Let c ∈ L1 (µ) := {f : X → R, |f (x)|µ(dx) < ∞}. Suppose that µ is an invariant probability
measure forPan X-valued Markov chain Xk . Then, by the individual ergodic theorem for µ almost everywhere x ∈ X:
T
limT →∞ T1 t=1 c(Xt ) = c(x)µ(dx), Px almost surely (that is conditioned on x0 = x, with probability one, the
R

above holds). Furthermore, again with c ∈ L1 (µ), for µ almost everywhere x ∈ X


XT  Z
1
lim Ex c(Xt ) = c(x)µ(dx), (7.47)
T →∞ T
t=1

On the other hand, the positive Harris recurrence property allows the almost sure convergence to take place for every
initial condition: If µ is the invariant probability measure for a positive Harris recurrent Markov chain, it follows that
for all x ∈ X and for every c ∈ L1 (µ)
T Z
1X
lim c(Xt ) = c(x)µ(dx), (7.48)
T →∞ T
t=1

almost surely. However, as discussed earlier, it is not generally true that


T Z
1 X
lim Ex [ c(Xt )] = c(x)µ(dx),
T →∞ T
t=1

for all x ∈ X. Thus, we can not in general relax the boundedness condition for the convergence of the expected costs.
With c bounded, for all x ∈ X
XT  Z
1
lim Ex c(Xt ) = c(x)µ(dx) (7.49)
T →∞ T
t=1

In the following, we follow the arguments in [16].

Theorem 7.4.1 a) Under (7.4.1) A, B, C, D there exists an optimal measure in G. b) Under (7.4.1) A’, B’, D, E, there
exists a policy in which is optimal for the control problem given in (7.64) for every initial condition.

Proof. a) Consider (7.4.1) A, B, C, D. By (B, C) we have that the set of policies γ which lead to a finite cost is so
γ γ
 ⟨µT , c⟩ < ∞ for all T , which implies that {µT , T > 0} is tight. Thus along some subsequence µtk → µ ∈
that
( (x, u) : x ∈ X, u ∈ U(x) ). As shown in the paragraph preceding (7.4.2), µ ∈ G.
Furthermore, under hypothesis (A), the set K = {(x, U(x)), x ∈ X} is closed by [164, Lemma D.3]. Thus, by the
Portmanteau theorem that every weak limit of a converging sequence of probability measures K) = 1 is also supported
on K.
Consider a sequence {µk }k∈N ⊂ G such that ⟨µk , c⟩ → δ ∗ as k → ∞, the sequence µtk is tight by inf-compactness,
and any limit point µ∗ of this sequence is in G with µ∗ ({(x, U(x)), x ∈ X}) = 1. Thus, by [164, Prop. D.8] we have
an optimal control policy ϕ. Taking limits as in (7.46), we obtain ⟨µ∗ , c⟩ = δ ∗ . This establishes the first part of the
theorem.
Define a stationary policy γ via the disintegration

µ∗ (dx, du) = γ∗ (du|x) π∗ (dx) (7.50)


160 7 The Average Cost Problem

µ∗ almost surely. Note that via this disintegration the control γ∗ is defined π∗ -a.e. Let ϕ ∈ be any policy that agrees
with γ∗ on the support of π∗ .

b) Under (A’, B’, D), via (7.45) and that T f ∈ Cb ( (x, u) : x ∈ X, u ∈ U(x) ) by Assumption 5.2.1(i)), we have
that G is compact; we also have that Portmanteau theorem applies as in part a). By hypothesis (E), since the chain
under an optimal ϕ, is Harris recurrent, π∗ is its unique invariant probability measure. Optimality of ϕ, for every initial
condition, then follows by positive Harris recurrence given that c is bounded via (7.49) under hypothesis (B’). Thus,
J(x, ϕ) = ⟨µ∗ , c⟩ and µ∗ is optimal.

Theorem 7.4.1 can be stated under weaker assumptions. See, for example, [13, Theorem 2.1] among other references
in the literature.
In general, in the absence of (7.4.1) (E), there is a question of reachability. Suppose that the chain under the policy ϕ as
defined in the proof of Theorem 7.4.1 is a T model (see [309]). Then, as asserted in [309, Theorem 6.1], the Doeblin
decomposition of the state space contains, in general, a countable  collection of maximal Harris sets. In particular,
we have a decomposition into the disjoint union X = ∪i∈N Hi ∪ E , where each Hi is a maximal Harris set with
invariant measure πi , and E is transient. Now, by part (ii) of Theorem 6.1 in [309], only a finite number of the sets
Hi may have a nonempty intersection with any given compact set. This implies that π∗ can always be expressed as
a convex combination of finitely many ergodic invariant measures. Thus, if the Markov Chain is not recurrent, the
stationary policy defined above, in general, is only optimal in a restricted set of initial conditions. On implications
related to insensitivity to such initial state dependence, the reader is referred to [213] and [165, Prop. 11.4.4(c) and
Lemma 11.4.5(a)], among other references, for further results on sample path average cost optimality and expected
average cost optimality. See [14] for further discussions.
Following the above, there exists an optimal expected empirical occupation measure, say v. This defines the optimal
stationary control policy by the decomposition:

dv(·, du)
µ(·|u) = R (u),
d u∈U
v(·, du)

v almost surely, where d· denotes the Radon-Nikodym derivative.

7.4.3 General state/action spaces under strong continuity in actions

There are many important applications where the kernel T is not weakly continuous. For example, consider dynamics
described by a stochastic difference equation on Rd of the form

Xn+1 = F (Xn , Un ) + Wn , n = 0, 1, 2, . . . ,

where X = Rn and the Wn ’s are independent and identically distributed (i.i.d.) random vectors whose distribution
has a bounded and continuous density function. We assume that F is bounded and u 7→ F (x, u) is continuous for all
x ∈ X. It is clear that the transition kernel T is not, in general, weakly continuous. However, it satisfies the following
hypothesis.

(H2)The transition kernel T satisfies the following:


R
(a) For any x ∈ X, the map u 7→ f (z)T (dz|x, u) is continuous for every bounded measurable function f . That
is, Assumption 5.2.2(i)) holds.
(b) There exists a finite measure ν majorizing T , that is

T (dy|x, u) ≤ ν(dy) , x ∈ X, u ∈ U . (7.51)


7.4 The Convex Analytic Approach to Average Cost Markov Decision Problems 161

If in addition, the distribution of Wn has a continuous, bounded, and a strictly positive probability density function
(a non-degenerate Gaussian distribution satisfies this condition), then positive Harris recurrence can be established
by Lebesgue-irreducibility and a uniform countable additivity condition for compact sets following [310, Condition
A], which leads to the presence of accessible compact petite sets (where one can take V (x) = x2 as the Lyapunov
function). For more details see [269, Example 3.1].

Assumption 7.4.2 The following hold:


(A) The state and action spaces X and U are Polish. The set K = {(x, U(x)), x ∈ X} is measurable (see [164, Lemma
D.3] for conditions) and the set-valued map U : X → B(U) is compact-valued.
(A’) The state and action spaces X and U are compact. The set K is measurable and set-valued map U : X → B(U) is
compact-valued.

(B) The non-negative running cost function c(x, u) is continuous in u ∈ U(x) for every x ∈ X and c : (x, u) : x ∈
X, u ∈ U(x) → R is inf-compact.
(B’) The cost function c is bounded, and continuous in u ∈ U(x) for every x ∈ X.
(C) There exists a policy and an initial state leading to a finite cost η ∈ R+ .
(D) (H2) holds.
(E) Under every stationary policy, the induced Markov chain is Harris recurrent.

We again recall the w-s topology studied in Appendix (Section D.4).


By Theorem D.4.1 (see [284, Theorem 3.10] or [27, Theorem 2.5]), (7.51) implies that setwise sequential pre-
compactness of marginal measures on the state ensures that every weakly converging sequence of mean empirical
occupation measures also converges in the w-s sense. (7.51) implies setwise sequential pre-compactness by [269,
Proposition 3.2], which in turn builds on [167, Corollary 1.4.5]; see also [152, Theorem 4.17].
First, note the following counterpart to Lemma 7.4.2.

Lemma 7.4.3 [16] Under (H2), the limit of any w-s converging subsequence of mean empirical occupation measures
is in G.

Proof. We follow the notation used in the


 discussion leading to Lemma 7.4.2. Suppose that, along some subsequence
{tk } ⊂ N, µγt converges to some µ ∈ ( (x, u) : x ∈ X, u ∈ U(x) ) in the w-s sense, which we denote as µγtk ⇒ µ.
As in (7.44) we have the triangle inequality

µ(f ) − µT (f ) ≤ µ(f ) − µγtk (f ) + µγtk (f ) − µγtk T (f )


(7.52)
+ µγtk T (f ) − µT (f )

for f ∈ Mb (X). If (H2) holds, the first term on the right hand side of (7.52) vanishes as k → ∞ by w-s convergence,
while the second term does so by (7.43). We have

µγtk T (f ) = µγtk (T f ), (7.53)

Since T f is continuous in u for every fixed x, by (H2), it follows that the third term also vanishes as k → ∞ by the
w-s convergence µγtk ⇒ µ. This shows that µ(A, U) = µT (A) for all A ∈ B(X), which implies that µ ∈ G. ⋄

Theorem 7.4.2 [16] a) Under (7.4.2) A, B, C, D, there exists an optimal measure in G. b) Under (7.4.2) A’, B’, D, E,
there exists a policy in which is optimal for the control problem given in (7.64) for every initial condition.
162 7 The Average Cost Problem

7.4.4 Optimality of deterministic stationary policies

In this section, we will present conditions on the optimality of deterministic policies via the convex analytic method;
see [16] for a detailed review and [231, Proposition 9.2.5], [59, Lemma 2.4] and [160, Corollary 5.4(b)] for related
results on the optimality of stationary and deterministic policies arrived at via different approaches.
Hernandez-Lermá [160, Theorem 5.3] shows that an average cost optimal randomized policy ϕ, with invariant measure
πϕ satisfies the ACOI πϕ almost everywhere:
Z
g + h(x) ≥ c(x, ϕ(x)) + h(x′ )T (dx′ |x, ϕ(x)) (7.54)

where h is bounded from below. If one can ensure that the above holds for all x ∈ X (and not just πϕ almost everywhere)
[160, Prop. 5.2] shows that under this condition on h, (7.54) implies that such a policy is indeed optimal. Again, if the
above holds for all x ∈ X, by utilizing Blackwell’s theorem 5.1.1 on optimality of deterministic policies we can replace
ϕ with a deterministic f ∈ ΓSD , which will then be optimal [160, Corollary 5.4(b)].
This approach can be generalized to the case where the induced Markov chain is not positive Harris recurrent, but when
the action space is finite [16]: Accordingly, one can relax the condition of (7.54) holding for every x. Let g be a constant
and h : X → R+ , f : X → P(U) be so that for all x ∈ B for some Borel set B ⊂ X,
 Z  Z  Z 
g + h(x) ≥ c(x, f (x) + h(x′ )T (dx′ |x, f (x)) := c(x, u) + h(x′ )T (dx′ |x, u) f (du|x) (7.55)

Lemma 7.4.4 Let (7.55) hold with


1 γ∗
lim inf E [h(Xn )] ≥ 0, (7.56)
n→∞ n x

for all x ∈ B where γ ∗ = {f, f, f, · · · } and P γ (x, B) = 1 for all x ∈ B. Then the stationary (possibly randomized)
policy γ ∗ = {f, f, f, · · · } satisfies
g ≥ J(x, γ ∗ ),
for all x ∈ B.

Proof. We have, as in the proof of Theorem 7.1.1,


Z
E γ [h(Xt )|x[0,t−1] , u[0,t−1] ] = h(y)P (Xt ∈ dy|xt−1 , ut−1 ) (7.57)
y
Z
= c(xt−1 , ut−1 ) + h(y)P (dy|xt−1 , ut−1 ) − c(xt−1 , ut−1 ) (7.58)
y

By iterated expectations,
n
X 
∗ ∗
Exγ h(Xt ) − E γ [h(Xt )|X[0,t−1] , U[0,t−1] ] = 0
t=1

Now, under γ we have that B is an absorbing set and thus, by (7.55) holding on the absorbing set, the following will
apply almost surely with X0 = x where x ∈ B:
Z
γ∗
E [h(Xt )|x[0,t−1] , u[0,t−1] ] = h(y)P (Xt ∈ dy|xt−1 , ut−1 ) (7.59)
y
Z
= c(xt−1 , f (xt−1 )) + h(y)P (dy|xt−1 , f (xt−1 )) − c(xt−1 , f (xt−1 )) (7.60)
y
≤ g + h(xt−1 ) − c(xt−1 , f (xt−1 )) (7.61)

Iterating the above and dividing by n, we arrive at


7.4 The Convex Analytic Approach to Average Cost Markov Decision Problems 163
n
1 γ∗ 1 γ∗ 1 γ∗ X
g − Ex [h(Xn )] + Ex [h(X0 )] ≥ Ex [ c(Xt−1 , Xt−1 )].
n n n t=1

Taking the limsup on both sides (and replacing lim sup with lim inf by reversing the negative sign on the left), and
(7.56) holding for γ ∗ = {f, f, f, · · · }, we establish the desired bound. ⋄
In particular if we have that g is a lower bound on the optimal cost (say via the convex analytic method), we can claim
that γ ∗ is optimal for all initializations X0 = x where x ∈ B. Now, the analysis in [160, Theorem 5.3] shows that if
we have an optimal invariant measure, then this leads to (7.54) for some randomized ϕ on a set of measure 1 under πϕ
with h bounded from below. Building on [48, 51], via [160, (5.7)], this implies the existence of a deterministic control
policy k which is defined on B and which satisfies
 Z 
′ ′
g + h(x) ≥ c(x, k(x)) + h(x )T (dx |x, k(x)) (7.62)

However, with κ∗ = {k, k, k, · · · }, to be able to claim the optimality of k over B via Lemma 7.1.3, we need to show

P κ (x, B) = 1 for all x ∈ B; that is an absorbing set under k should be a subset of the absorbing set under ϕ when
X0 = x with x ∈ B. If the induced Markov chain under ϕ is positive Harris recurrent, then [160, Theorem 5.3(b)]
shows that 7.54 holds everywhere (that is, for all x ∈ X), and the result follows. Additionally, when U is countable,
this result also follows via the following argument: By Blackwell’s theorem 5.1.1 and by the measurable selection
theorem of RBlackwell and Ryll-Nardzewski [51],Rk can be (without loss) constructed such that for all x : k(x) ∈ {u :
(c(x, u) + h(x′ )T (dx′ |x, u)) ≤ c(x, ϕ(x)) + h(x′ )T (dx′ |x, ϕ(x))} ∩ {u : ϕ(u|x) > 0}. In this case, it follows
by expressing the transition probabilities in terms of the countable collection of control realizations, we will have that

P κ (x, B) = 1 for all x ∈ B. This leads to the following result.

Theorem 7.4.3 Assume that either Theorem 7.4.1 or Theorem 7.4.2 apply. Let µ∗ be an optimal invariant measure.
Define a stationary policy γ via the disintegration

µ∗ (dx, du) = γ∗ (du|x) π∗ (dx) (7.63)

µ∗ almost surely. Take ϕ ∈ ΓS be any policy that agrees with γ∗ on the support of π∗ .
(i) [160] If the induced Markov chain under ϕ is positive Harris recurrent, then the optimal policy can be assumed
deterministic.
(ii) [16] If the induced Markov chain under an optimal policy is not positive Harris recurrent, then with U countable,
on the support of π∗ , ϕ can be assumed to be deterministic. This would lead to an optimal policy for all initial
states x with X0 = x where x ∈ supp (π∗ ).

There exist alternative arguments in the literature: [231, Proposition 9.2.5], [59, Lemma 2.4] focus on the properties
of return sets for countable state/action space models and characterize conditions under which an optimal policy is
deterministic, and the analysis in [59, Section 3.2] builds on Schauder’s fixed point theorem under restrictive regularity
conditions on a continuous space model. Furthermore, it can be shown that, under mild ergodicity conditions, deter-
ministic policies are dense in the sense that the performance under stationary deterministic policies is dense in the set
of performance values under randomized stationary policies [16].

Remark 7.4 (ACOI through duality with the convex analytic method). Though not directly related to the discussion
above, one could note that there is a further duality relationship between ACOI (see Definition 7.1.2) and the convex
analytic method, see [165, Chapter 12, p. 221].

7.4.5 Sample-path optimality

The above optimality arguments also apply in the somewhat stronger sample-path sense, rather than only in expectation.
164 7 The Average Cost Problem

Finite state/action setup

Consider the following:


T
1X
inf lim sup [c(xt , ut )], (7.64)
γ∈ΓA T →∞ T t=1

where there is no expectation. The above is known as the sample path cost. Let the state and action spaces be finite. Let
Ft be the σ−field generated by {xs , us , s ≤ t}, under any given admissible policy. Define a Ft measurable process
with A ∈ B(X):
X t X 
Ft (A) = 1{xs ∈A} − t P (A|x, u)vt (x, u) ,
s=1 X×U

where
t−1
X
vt (x, u) = 1{xs =x,ut =u} ,
s=0

is the empirical occupation measure. Note that the above can also be written as
t−1 
X X 
Ft (A) = 1{xs+1 ∈A} − P (A|x, u)1{xs =x,ut =u}
s=0 X×U

Thus, for t ≥ 1,

E[Ft (A)|Ft−1 ]
X t t−1 X
X 
=E 1{xs ∈A} − P (xs+1 ∈ A|xs = x, us = u)1{(xs ,us )=(x,u)} |Ft−1
s=1 s=0 X×U
 X  
=E 1{xt ∈A} − P (xs+1 ∈ A|xs = x, us = u)1{(xt−1 ,ut−1 )=(x,u)} |Ft−1
X×U
t−1
X t−2 X
X 
+ 1{xs ∈A} − P (xs+1 ∈ A|xs = x, us = u)1{(xs ,us )=(x,u)}
s=1 s=0 X×U
=0 (7.65)
t−1
X t−2 X
X  
+ 1{xs ∈A} − P (xs+1 ∈ A|xs = x, us = u)1{(xs ,us )=(x,u)} |Ft−1 (7.66)
s=1 s=0 X×U
= Ft−1 (A), (7.67)

where (7.65) follows from the fact that E[1{xt ∈A} |Ft−1 ] = P (xt ∈ A|Ft−1 ).
We have then that
E[Ft (A)|Ft−1 ] = Ft−1 (A) ∀t ≥ 0,
and {Ft (A)} is a martingale sequence.
Furthermore, Ft (A) is a bounded-increment martingale since |Ft (A) − Ft−1 (A)| ≤ 1. Hence, for every T > 2,
{F1 (A), . . . , FT (A)} forms a martingale sequence with uniformly bounded increments, and we could invoke the
Azuma-Hoeffding inequality [97] to show that for all x > 0

Ft (A) 2
P (| | ≥ x) ≤ 2e−2x t
t
Finally, invoking the Borel-Cantelli Lemma (see Theorem B.2.1) for the summability of the estimate above, that is:
7.4 The Convex Analytic Approach to Average Cost Markov Decision Problems 165

X 2
2e−2x t
< ∞, ∀x > 0,
n=1

we deduce that
Ft (A)
lim = 0 a.s.
t→∞ t

Thus,  
X
lim vT (A) − P (A|x, u)vT (x, u) = 0, A⊂X
T →∞
X×U

Thus, somewhat similar to the arguments in (7.39), every converging subsequence would have to be in the set G defined
in (7.40).
P
Let ⟨v, c⟩ := v(x, u)c(x, u). Now, we have that

lim inf ⟨vT , c⟩ ≥ δ ∗


T →∞

since for any sequence vTk which converges to the liminf value, there exists a further subsequence vTk′ (due to the
(weak) compactness of the space of occupation measures) which has a weak limit, and this weak limit is in G. Then,

lim ⟨vTk , c⟩ = ⟨ lim



vTk , c⟩ ≥ γ ∗ .
Tk →∞ Tk →∞

Furthermore, this cost is attained by an optimal stationary policy as a consequence of positive Harris recurrence.
Note that, the above would lead to the same for the average cost problem as well (though as we studied earlier, a more
direct argument is applicable for the average cost setup):

lim inf E[⟨vT , c⟩] ≥ E[lim inf ⟨vT , c⟩] ≥ γ ∗ .


T →∞ T →∞

The standard Borel setup

As we observed, the discussion in Section 7.4.1 applies to the sample path optimality also. We now discuss a more
general setting where the state and action spaces are Polish. Let ϕ : X → R be a continuous and bounded function.
Define:
T
1X
vT (ϕ) = ϕ(x, u).
T t=1

Define a Ft measurable process, with π an admissible control policy (not necessarily stationary or Markov):
t
X  Z Z 
Ft (ϕ) = ϕ(xt ) − t ϕ(x′t )P π (dx′t , du′t )|x)vt (dx)
s=1 P×U

(7.68)

As earlier, we define GX to be the following set in this case.


Z
GX = {η ∈ P(X × U) : η(D) = P (D|z)η(dz), ∀D ∈ B(X)}.
X×U

Consider the following sample-path cost:


166 7 The Average Cost Problem
T
1X
inf lim sup [c(xt , ut )], (7.69)
γ T →∞ T t=1
P
where there is no expectation. Let ⟨v, c⟩ := v(x, u)c(x, u). If one can guarantee that every sequence of empirical
measures {vt } would have a converging subsequence to some measure v, we would have that

lim ⟨vTk , c⟩ ≥ ⟨ lim



vTk , c⟩,
Tk →∞ Tk →∞

by the fact that for c continuous, non-negative if vk → v,

lim inf ⟨vk , c⟩ ≥ ⟨v, c⟩.


k→∞

Since for any sequence vTk which converges to the liminf value, there exists a further subsequence vTk′ (due to the
(weak) compactness of the space of occupation measures) which has a weak limit, and this weak limit is in G. By
Fatou’s Lemma:
lim ⟨vTk , c⟩ = ⟨ lim

vTk , c⟩ ≥ γ ∗ .
Tk →∞ Tk →∞

To apply the convex analytic approach, we require that under any admissible policy, the set of sample path occupation
measures would be tight, for almost every sample path realization. If this can be established, then the result goes through
not only for the expected cost, but also the sample-path average cost, as discussed for the finite state-action setup.
Researchers in the literature have tried to establish conditions which would ensure that the set of empirical occupa-
tional measures are tight. These typically follow one of two conditions: Either cost functions are near-monotone type
conditions [59] (this includes, but is more general than, the condition: lim|x|→∞ inf u∈U c(x, u) = ∞) or behave like
moments [213] (when X × U is locally compact, there exists a sequence of compact sets Kn so that X × U = ∪n Kn
/ n c(x, u) = ∞), or the Markov chain satisfies strong recurrence properties [59] [15, Chapter 3].
withlimKn ↑X inf (x,u)∈K
Under such conditions, the sequence of empirical occupation measures {vn } which give rise to a finite cost are almost
surely tight, every such sequence has a convergent subsequence and thus the arguments above apply: Every expected
average-cost optimal policy is also sample-path optimal provided that the initial condition belongs to the support of the
invariant probability measure under an optimal policy.
Note also that one often can obtain more relaxed conditions for sample path optimality when compared with expected
cost optimality as a consequence of the ergodic theorems for positive Harris recurrent Markov chains (Section 3.4.1).
Note, however, that for sample path optimality, we need to invoke weak continuity almost surely, by the martingale
argument above, and accordingly the measurable selection criteria will need to be under Assumption 5.2.1.

7.5 Constrained Markov Decision Processes

Consider the following average cost problem:


T −1
1 γX
inf J(x, γ) = inf lim sup Ex c(xt , ut ) (7.70)
γ γ T →∞ T t=0

subject to the constraints:


T −1
1 γX
lim sup Ex di (xt , ut ) ≤ Di (7.71)
T →∞ T t=0

for i = 1, 2, · · · , m where m ∈ N.
A linear programming formulation leads to the following result.
7.7 Exercises 167

Theorem 7.5.1 [265] [7] Let X, U be countable. Consider (7.70-7.71). An optimal policy will randomize between at
most m + 1 deterministic policies.

Ross also discusses a setup with one constraint where a non-stationary history-dependent policy may be used instead
of randomized stationary policies.
Finally, the theory of constrained Markov Decision Processes is also applicable to Polish state and action spaces, but
this requires further technicalities. If there is an accessible atom (or an artificial atom as considered earlier in Chapter
3) under any of the policies considered, then the randomizations can be made at the atom.

7.6 Bibliographic Notes

7.7 Exercises

Exercise 7.7.1 Let X, U be finite sets and consider the occupation measure corresponding to a controlled Markov chain
under some arbitrary admissible control policu:
T −1
1 X
vT (A × B) = 1{(xt ,ut )∈A×B} , A ⊂ X, B ⊂ U.
T t=0

While proving that the limit of such a measure process lives in a specific set, the following is used, which you are asked
to prove. Let γ be some arbitrary but admissible control policy and let Ft be the σ−field generated by {xs , us , s ≤ t}.
Define a Ft measurable process
t
X X 
Ft (A) = 1{xs ∈A} − t P (A|x)vt (x, u) ,
s=1 X×U

Show that, {Ft (A), t ∈ Z+ } is a martingale sequence.


Hint: Observe that for all t ∈ {1, 2, . . . , T }
t
X X 
1{xs ∈A} − t P (x1 ∈ A|x0 = x, u0 = u)vt (x, u)
s=1 X×U
t
X t−1 X
X 
= 1{xs ∈A} − P (xs+1 ∈ A|xs = x, us = u)1{(xs ,us )=(x,u)} (7.72)
s=1 s=0 X×U

Then, show that E[Ft (A)|Ft−1 ] = Ft−1 (A). You may follow Exercise 4.5.11.

Exercise 7.7.2 a) Let, for a Markov control problem, xt ∈ X, ut ∈ U, where X and U are finite sets denoting the state
space and the action space, respectively. Consider the optimal control problem of the minimization of
T −1
1 γ X
lim sup Ex [ c(xt , ut )],
T →∞ T t=0

where c is a bounded function. Further assume that under any stationary control policy, the state transition kernel
P (xt+1 |xt , ut ) leads to an irreducible Markov Chain.
Does there exist an optimal control policy? Propose a method to find an optimal policy.
b) Is the optimal policy also sample-path optimal?
168 7 The Average Cost Problem

Exercise 7.7.3 Consider a controlled Markov Chain with state space X = {0, 1}, action space U = {0, 1}, and
transition kernel for t ∈ Z+ :
P (xt+1 = 1|xt = 0, ut = 1) = α ∈ (0, 1)
P (xt+1 = 1|xt = 0, ut = 0) = β ∈ (0, 1)
1
P (xt+1 = 1|xt = 1, ut = 0) = P (xt+1 = 1|xt = 1, ut = 1) =
2
Let
c(0, 1) = κ ∈ R+ , c(0, 0) = 1
c(1, 0) = c(1, 1) = 1

Suppose, the goal is to minimize the quantity


T −1
1 γ X
lim sup E0 [ c(xt , ut )],
T →∞ T t=0

over all admissible policies γ ∈ ΓA .


Find the optimal policy and the optimal cost, as a function of α, β, κ. Explain your answer and how you arrived at your
solution.

Exercise 7.7.4 Consider a controlled Markov chain with state space X = {0, 1}, action space U = {0, 1}, and
transition kernel for t ∈ Z+ :
P (xt+1 = 1|xt = 0, ut = 1) = 1
1
P (xt+1 = 1|xt = 0, ut = 0) =
2
1
P (xt+1 = 1|xt = 1, ut = 0) = P (xt+1 = 1|xt = 1, ut = 1) = .
2
Let a cost function c : X × U → R+ be given by

c(0, 1) = κ ∈ R+ , c(0, 0) = 1
1
c(1, 0) = , c(1, 1) = 1.
2

Suppose that the goal is to minimize the quantity


T −1
1 γ X
lim sup E0 [ c(xt , ut )],
T →∞ T t=0

over all admissible policies γ ∈ γA . Recall that a policy is admissible if the controller has access to {xs , s ≤ t; ul , l ≤
t − 1} at time t ∈ Z+ .
Find an optimal policy and the optimal expected cost explicitly, as a function of κ. Explain your answer and how you
arrived at your solution.

Exercise 7.7.5 Consider a two-state, controlled Markov Chain with state space X = {0, 1}, and transition kernel for
t ∈ Z+ :
P (xt+1 = 0|xt = 0) = u0t
P (xt+1 = 1|xt = 0) = 1 − u0t
P (xt+1 = 1|xt = 1) = u1t
P (xt+1 = 0|xt = 1) = 1 − u1t .
7.7 Exercises 169

Here u0t ∈ [0.2, 1] and u1t ∈ [0, 1] are the control variables. Suppose, the goal is to minimize the quantity
T −1
1 γ X
lim sup E0 [ c(xt , ut )],
T →∞ T t=0

where
c(0, u0 ) = 1 + u0 ,
c(1, u1 ) = 1.5, ∀u1 ∈ [0, 1],
with given α, β ∈ R+ .
Find an optimal policy and find the optimal cost.
Hint: Consider deterministic and stationary policies and analyze the costs corresponding to such policies.

Exercise 7.7.6 [Machine repair revisited] Recall Exercise 7.7.6 with an average cost formulation. Show that there
exists an optimal control policy and that this policy is stationary.

Exercise 7.7.7 For the model considered in Section 7.2, under Assumption 7.2.2, establish that Tz ) is a contraction on
the space of bounded functions h with h(z) = 0, with modulus α.

Exercise 7.7.8 (Risk-sensitive average cost criterion) Let X, U be finite and consider the following risk-sensitive cri-
terion:   PN −1 
∗ 1 γ c(xm ,um )
λ = inf lim sup log E e m=0
γ∈ΓA N →∞ N

For this criterion, show that the verification theorem (dynamic programming equation) satisfies [67]:
 X 
λ∗ V (x) = min ec(x,u) T (x1 |x0 = x, u0 = u)V (x1 )
u∈U
x1

If one views the operator on the right as a function of V , we have an eigenvalue problem with λ∗ being the optimal
cost. Hint. Divide both sides by λ∗ , and apply iteratively the inequality under an arbitrary admissible control. Show
that equality holds when an optimal policy is considered.

Exercise 7.7.9 Study [266] as an example where an average cost optimal policy may not be stationary (and even
possibly randomized stationary).

Exercise 7.7.10 In some problems one needs to relate the discounted cost problem, a finite horizon average cost prob-
lem and an infinite horizon average cost problem. Under what conditions do we have that

lim (1 − β)Jβ (x0 ) → J ∗ (x0 )?,


β→1

and with J T as given in (7.9),


lim inf J T (x0 , γ) → J ∗ (x0 )?
T →∞ γ∈ΓA

Exercise 7.7.11 a) For an infinite horizon discounted cost partially observed Markov decision problem with finite state,
action and measurement spaces, suppose that we wish to restrict the policies to be stationary control policies which
only are based on the most recent observation; that is ut = γ(yt ) for some γ : Y → U (clearly, this is suboptimal
among all admissible policies, as the analysis in the Chapter shows). Given this restrictive class of policies, can one
obtain an optimal policy through linear programming? b) Can you consider a setup where an optimal policy above
may not be optimal among all policies (e.g., an optimal one-memory policy may not be stationary)? Hint: Consider
linear systems theory (stationary output feedback vs. time-varying output feedback).
8

Numerical and Approximation Methods

In this chapter, we will first study several computational algorithms, first in the context of finite space MDPs. After
this, we will study rigorous approximation methods for continuous (standard Borel) space MDPs and POMDPs.

8.1 Value and Policy Iteration Algorithms

8.1.1 Value Iteration

Consider expected discounted cost criterion, for some β ∈ (0, 1),



X
Jβ (x0 , γ) = Exγ0 [ β t c(xt , ut )], (8.1)
t=0

to be minimized. The Value Iteration Algorithm was presented earlier in Theorems 5.5.1 and 5.5.2; the algorithm for
the bounded setup is re-stated in the following.

Theorem 8.1.1 Suppose the cost function c is bounded, non-negative, and one of the measurable selection conditions
(Assumption 5.2.1 or Assumption 5.2.2) applies. Then, there exists a unique solution to the discounted cost optimality
equation  Z 
v(x) = min c(x, u) + β v(y)T (dy|x, u) , x∈X
u∈U X

Furthermore, the optimal cost (value function) is obtained by a successive iteration of policies (known as the Value
Iteration Algorithm):
 Z 
vn (x) = min c(x, u) + β vn−1 (y)T (dy|x, u) , ∀x, n ∈ N (8.2)
u∈U X

For any v0 ∈ L∞ (X) (or Cb (X) under measurable selection Condition 1 in Assumption 5.2.1), the sequence converges
to a unique fixed point. If v0 (x) = 0 for all x ∈ X, then vn (x) ↑ v(x) for all x ∈ X. Under the measurable selection
Assumption 5.2.1, the limit is also continuous with v0 ∈ Cb (X).

We also recall that under Assumptions 5.2.1 and 8.2.5, the value function is Lipschitz; see Theorem 5.5.3.

8.1.2 Policy Iteration

We now discuss the Policy Iteration Algorithm. Let X be countable and c be a bounded cost function. Consider again
(8.1). Let γ0 := {γ0 , γ0 , γ0 , · · · , γ0 , · · · } ∈ ΓS denote a deterministic stationary policy (which naturally leads to a
172 8 Numerical and Approximation Methods

finite discounted expected cost here). Let this expected cost be W0 (·); that is,

X
W0 (x) = Exγ0 [ β k c(xk , γ0 (xk ))], x∈X
k=0

Then, similar to (5.23), and using the stationarity of γ0 , we obtain


X
W0 (x) = c(x, γ0 (x)) + β W0 (x′ )P (xt+1 = x′ |xt = x, ut = γ0 (x))
x′

Let  
X
′ ′
T(W0 )(x) = min c(x, u) + β W0 (x )P (xt+1 = x |xt = x, ut = u) .
u∈U
x′

Clearly T(W0 ) ≤ W0 pointwise (in x). Now, let γ1 be such that


X
T(W0 )(x) = c(x, γ1 (x)) + β W0 (x′ )P (x1 = x′ |x0 = x, u0 = γ1 (x)) (8.3)
x′

Observe that, iterative application of (8.3) one more time (to W0 (x′ ) above with T(W0 )(x′ ) ≤ W0 (x′ )) leads to
1
X 
W0 (x) ≥ Exγ1 β k c(xk , γ1 (xk )) + β 2 W0 (x2 ) .
k=0

and, continuing further, for any x ∈ X, n ∈ Z+ , we arrive at


n−1
X
γ1
W0 (x) ≥ Ex [ β k c(xk , γ1 (xk )) + β n W0 (xn )]. (8.4)
k=0

Taking the limit n → ∞, this leads to the relation

W0 (x) ≥ T(W0 )(x) ≥ W1 (x),

with

X
W1 (x) = Exγ1 [ β k c(xk , γ1 (xk ))]
k=0

so that X
W1 (x) := c(x, γ1 (x)) + β W1 (x′ )P (xt+1 = x′ |xt = x, ut = γ1 (x)).
x′

We can interpret the steps of the discussion above as follows: We start with the policy, {γ0 , γ0 , γ0 , · · · } and then make
the point that the policy {γ1 , γ0 , γ0 , · · · } is a better one and then {γ1 , γ1 , γ0 , · · · } is a better one and ultimately the
policy {γ1 , γ1 , γ1 , · · · , γ1 , · · · } is a better policy than what we started with.
We then continue this procedure for m = 2 by replacing W1 with W0 above, to arrive at
 X 
′ ′
T(W1 )(x) = min c(x, u) + β W1 (x )P (xt+1 = x |xt = x, ut = u)
u∈U
x′
X
= c(x, γ2 (x)) + β W1 (x′ )P (x1 = x′ |x0 = x, u0 = γ2 (x)) (8.5)
x′

and ultimately W2 (x) ≤ T(W1 )(x) ≤ W1 (x), where


8.1 Value and Policy Iteration Algorithms 173

X
W2 (x) = Exγ2 [ β k c(xk , γ1 (xk ))]
k=0

Then, we repeat the process for m > 2 with


 X 
′ ′
T(Wm )(x) = min c(x, u) + β Wm (x )P (xt+1 = x |xt = x, ut = u)
u∈U
x′
X
= c(x, γm+1 (x)) + β Wm (x′ )P (x1 = x′ |x0 = x, u0 = γm+1 (x)) (8.6)
x′

and ultimately Wm+1 (x) ≤ T(Wm )(x) ≤ Wm (x), where



X
Wm+1 (x) = Exγm+1 [ β k c(xk , γm+1 (xk ))]
k=0

Remark 8.1. Note that for the case where X, U are finite, the following holds

X X
W1 (x) = Exγ1 [ β k c(xk , γ1 (xk ))] = c(x, γ1 (x)) + β W1 (xt+1 )T (dxt+1 |xt = x, ut = γ1 (x))]
k=0

More generally, for a given stationary policy γ:



X X
Jβ (x, γ) = Exγ [ β k c(xk , γ(xk ))] = c(x, γ(x)) + β Jβ (x′ , γ)P (xt+1 = x′ |xt = x, ut = γ(x)) (8.7)
k=0 x′

can be computed by solving the following matrix equation

W = cγ + βP γ W,

leading to
W = (I − βP γ )−1 cγ ,
where W is a column vector consisting of {W (x), x ∈ X}; cγ is a column vector consisting of elements {c(x, γ(x)), x ∈
X}; and P γ is a stochastic matrix with entries P γ (x, x′ ) = P (xt+1 = x′ |xt = x, ut = γ(x)) (note that (I − βP γ ) is
always invertible for β ∈ (0, 1)). Thus, the implementation of the policy iteration algorithm is quite simple.

Theorem 8.1.2 Through the policy iteration algorithm, there exists W : X → R such that Wn ↓ W pointwise in x,
provided that for some n ∈ N, Wn (x) < ∞ for x ∈ X. If γ = {f, f, · · · } satisfies

X
Exγ [ β k c(xk , f (xk ))] = W (x), (8.8)
k=0

then γ is optimal among all stationary policies. And if (5.29) holds with W replacing v, γ is optimal among all
policies (note that this always holds if c is bounded). For a problem with finite state and action spaces, convergence is
guaranteed in a finite number of stages and the resulting policy is optimal.

Proof. By (8.4) the sequence Wn ≥ T(Wn ) ≥ Wn+1 , and thus there is a limit W (since the cost per state is bounded
from below) and the limit satisfies W = T(W ), which is precisely the optimality equation (5.28). Since such a W
leads to a lower bound under any stationary policy by the construction of the algorithm, an argument similar to the one
in the proof of Lemma 5.5.4 leads to the result. Note that for finite models we have the condition, as noted in Lemma
5.5.4,
lim β t Exγ [W (xt )] = 0,
t→∞

by (8.8), since W is bounded. ⋄


174 8 Numerical and Approximation Methods

8.1.3 Receding Horizon Algorithms / Model Predictive Control

Roll-out algorithms, also known as sliding-horizon or receding horizon algorithms, are practically important. Such
an algorithm is provably near-optimal as the horizon length increases under some conditions. We refer the reader
to [163], [78], [96] and [40] among many other papers in this direction.

8.2 Approximation through Quantization of the State and the Action Spaces

For cases where the spaces are not finite or countable, numerical methods require approximation procedures. In the fol-
lowing, let X, U be standard Borel spaces with U(x) = U compact. We will arrive at rigorously justified approximation
algorithms and resulting convergence results, both via analytical as well as learning theoretic methods.

8.2.1 Finite Action Approximation to MDPs

Definition 8.2. A measurable function q : X → U is called a quantizer from X to U if the range of q, i.e., q(X) =
{q(x) ∈ U : x ∈ X}, is finite.

The elements of q(X) (the possible values of q) are called the levels of q.
Finite Action Approximate MDP: Quantization of the Action Space
Let dU denote the metric on U. Since the action space U is compact and thus totally bounded, one can find a sequence
of finite sets Λn = {an,1 , . . . , an,kn } ⊂ U such that for all n,

1
min dU (a, an,i ) < for all a ∈ U.
i∈{1,...,kn } n

In other words, Λn is a 1/n-net in U. In the rest of Section 8.2.1, we assume that the sequence {Λn }n≥1 is fixed. To
ease the notation in the sequel, let us define the mapping

Υn (f )(x) := arg min dU (f (x), a), (8.9)


a∈Λn

where ties are broken so that Υn (f )(x) is measurable.


Our main objective in this section is to find conditions on the components of the MDP under which there exists a
sequence of finite subsets {Λn }n≥1 of U for which the following holds:

(P) If for each n, MDPn is defined as the Markov decision process having the components X, Λn , p, c , then we
would like to find conditions under which the value function of MDPn converges to the value function of the original
MDP as n → ∞.

Near optimality of quantized policies under weak continuity

Consider (P) for MDPs with weakly continuous transition probability.


Recall Assumption 5.2.1, essentially repeated for convenience of the reader:

Assumption 8.2.1 (a) The one stage cost function c is bounded and continuous.
(b) The stochastic kernel T ( · |x, a) is weakly continuous in (x, a) ∈ X × U.
8.2 Approximation through Quantization of the State and the Action Spaces 175

(c) U is compact.

For any real-valued measurable function u on X, let T be given by


 Z 
(Tv)(x) := min c(x, u) + β v(y)T (dy|x, u) . (8.10)
u∈U X

Recall that here, T is the Bellman discounted cost optimality operator for the MDP considered earlier in (5.28). Anal-
ogously, let us define the Bellman optimality operator Tn of MDPn as
 Z 
Tn v(x) := min c(x, u) + β v(y)T (dy|x, u) . (8.11)
u∈Λn X

We have seen that both T and Tn are contraction operators. Furthermore, value functions of MDP and MDPn are
fixed points of these operators; that is, TJ ∗ = J ∗ and Tn Jn∗ = Jn∗ . Let us define v 0 = vn0 ≡ 0, and v t+1 = Tv t
and vnt+1 = Tn vnt for t ≥ 1; that is, {v t }t≥1 and {vnt }t≥1 are successive approximations to the discounted value
functions of the MDP and MDPn , respectively (via value iteration). The following can be shown, inductively for each
t = 0, 1, 2, · · · :

Lemma 8.2.1 [270, Lemma 3.19] Under Assumption 8.2.1, for any compact K ⊂ X and for any t ≥ 1, we have

lim sup |vnt (x) − v t (x)| = 0. (8.12)


n→∞ x∈K

The following theorem states that the optimal value function of MDPn converges to the optimal value function of the
original MDP. It can be proved by using Lemma 8.2.1 and taking into account that {v t }t≥1 and {vnt }t≥1 are successive
approximations to the value functions Jβ∗ and Jβ,n

, respectively.

Theorem 8.2.1 [273] [270, Theorem 3.16] Under Assumption 8.2.1, for any compact K ⊂ X, we have

lim sup |Jβ,n (x) − Jβ∗ (x)| = 0. (8.13)
n→∞ x∈K

The proof follows from a successive approximation argument applied iteratively for value iteration updates; see the
proof of Theorem 12.3.2 for an explicit analysis and Section 12.5.2 for further relations between the problem considered
here and the robustness problem considered there, where the approximation problem here is viewed as a particular
instance of robustness.
We state an approximation result analogous to Theorem 8.2.1 for the average cost criterion. For average cost criteria,
we impose relatively stronger ergodicity/minorization conditions on the controlled Markov chain. Note that this was
utilized to establish the existence of a solution to the average cost optimality equation in Section 7.2 (see Assumption
7.2.2).
Suppose that Assumption 5.2.1 and Assumption 7.2.2 hold. This implies that, as we have seen earlier in Section 7.3,
there is a solution to the average cost optimality equation (ACOE) and the stationary policy which minimizes this
ACOE is an optimal policy.

Theorem 8.2.2 [Average Cost] [273], [270, Theorem 3.22] Under Assumptions 5.2.1 and 7.2.2, the value functions
(that is, the optimal expected average cost) satisfy

lim |Vn∗ − V ∗ | = 0,
n→∞

where V ∗ and Vn∗ (n ≥ 1) (the value functions of the true model and the approximate model sequence, respectively)
do not depend on x.
176 8 Numerical and Approximation Methods

Remark 8.3. As we have observed earlier, when one considers partially observed MDPs (POMDPs), any POMDP can
be reduced to a (completely observable) MDP whose states are the posterior state distributions or beliefs of the observer.
Thus the results in this section are applicable to POMDPs as we will study in Section 9.2.1.

Near optimality of quantized policies under strong continuity in actions for each state

Consider the problem (P) for MDPs with strongly continuous transition probabilities in actions. Recall Assumption
5.2.2, repeated for reader’s convenience:

Assumption 8.2.2
(a) The one stage cost function c is nonnegative and bounded satisfying c(x, · ) ∈ Cb (U) for all x ∈ X.
(b) The stochastic kernel T ( · |x, u) is setwise continuous in u ∈ U for every x ∈ X.
(c) U is compact.

The following theorem states that for any f ∈ F, the discounted cost function of Υn (f ) converges to the discounted
cost function of f as n → ∞. Therefore, it implies that the discounted value function of the MDPn converges to the
discounted value function of the original MDP.

Theorem 8.2.3 [275][Discounted Cost] Consider problem (P) for the discounted cost under Assumption 8.2.2. Then,
for any stationary policy defined withf : X → U, J(Υn (f ), x) → J(f, x) as n → ∞, for all x ∈ X.

Observe that any deterministic stationary policy f defines a stochastic kernel on X given X via

Qf ( · |x) := T ( · |x, f (x)). (8.14)

Let Qtf denote the t-step transition probability of this Markov chain. If Qf admits a unique invariant probability
measure νf , then by [167, Theorem 2.3.4 and Proposition 2.4.2], there exists an invariant set Mf ∈ B(X) with full νf
measure such that for all x in that set we have
Z
V (f, x) = c(x, f (x))νf (dx). (8.15)
X

Assumption 8.2.3 Suppose Assumption 8.2.2 holds. In addition, we have


(d) For any f ∈ F, Qf has a unique invariant probability measure νf .
(e1) The set of invariant probability measures ΓF := {ν ∈ P(X) : νQf = ν for some f ∈ F} is relatively sequentially
compact in the setwise topology.
(e2) There exists x ∈ X such that for all B ∈ B(X), Qtf (B|x) → νf (B) uniformly in f ∈ F.
T
(f) M := f ∈F Mf ̸= ∅.

We note that the condition (e2) above may be slightly relaxed [347].
The following theorem states that for any f ∈ F, the average cost function of Υn (f ) ∈ Q(Λn ) converges to the average
cost function of f as n → ∞. In particular, the average value function of MDPn converges to the average value function
of the original MDP.

Theorem 8.2.4 [275][Average Cost] Let x ∈ M and f ∈ F. Then, we have V (Υn (f ), x) → V (f, x) as n → ∞, under
Assumption 8.2.3 with either (e1) or (e2).

One can also obtain rates of convergence results [275] [270].


8.2 Approximation through Quantization of the State and the Action Spaces 177

8.2.2 Finite State Approximation to MDPs

In this section we study the finite-state approximation problem for MDPs, by reducing them to finite state MDPs
obtained through quantization of the state space on a finite grid following [275] [270, Chapter 4]. Here two questions
could be posed:

(i) Q1 Under what conditions on the components of the MDP do the true cost functions of the policies obtained from
finite models converge to the optimal value function as the number of grid points goes to infinity?
(ii) Q2 Can we obtain bounds on the performance loss due to discretization in terms of the number of grid points if we
strengthen the conditions sufficient in (Q1)?

We will not discuss Q2 here, but address Q1. For Q2, the reader is referred to [270, 275]. We note that, under fur-
ther explicit regularity conditions, one can indeed arrive at rates of convergence to optimality. The approach to solve
problem (Q1) can be summarized as follows: First, we obtain approximation results for the compact-state case. We
find conditions under which a compact representation leads to near optimality for non-compact state MDPs: solve the
approximate MDP, and apply the optimal solution for the approximate MDP to the original MDP. We then obtain the
convergence of the finite-state models to non-compact models. Consider (Q1) for MDPs with compact state space.
We now establish near
S optimality under T finite state approximations. We start by choosing a collection of disjoint sets
{Bi }M
i=1 such that B
i i = X, and B i Bj = ∅ for any i ̸= j. Furthermore, we choose a representative state, yi ∈ Bi ,
for each disjoint set. For this setting, we denote the new finite state space by Y := {y1 , . . . , yM }, and the mapping
from the original state space X to the finite set Y is done via

q(x) = yi if x ∈ Bi . (8.16)

Furthermore, we choose a weighting measure π ∗ ∈ P(X) on X such that π ∗ (Bi ) > 0 for all Bi . We now define
normalized measures using the weight measure on each separate quantization bin Bi as follows:

π ∗ (A)
π̂y∗i (A) := , ∀A ⊂ Bi , ∀i ∈ {1, . . . , M }, (8.17)
π ∗ (Bi )

that is, π̂y∗i is the normalized weight measure on the set Bi , where yi belongs to.
We now define the stage-wise cost and transition kernel for the MDP with this finite state space Y using the normalized
weight measures. Indeed, for any yi , yj ∈ Y and u ∈ U, the stage-wise cost and the transition kernel for the finite-state
model are defined as
Z

C (yi , u) = c(x, u) π̂y∗i (dx),
Bi
Z

P (yj |yi , u) = T (Bj |x, u) π̂y∗i (dx). (8.18)
Bi

Having defined the finite state space Y, the cost function C ∗ and the transition kernel P ∗ , we can now introduce the
optimal value function for this finite model. We denote the optimal value function which is defined on Y by Jˆβ : Y → R.
Note that Jˆβ satisfies the following DCOE for any y ∈ Y
( )
X
Jˆβ (y) = inf C ∗ (y, u) + β Jˆβ (z)P ∗ (z|y, u) . (8.19)
u∈U
z∈Y

Note that we can easily extend this function over the original state space X by making it constant over the quantization
bins. In other words, if y ∈ Bi , then for any x ∈ Bi , we write

Jˆβ (x) := Jˆβ (y).

We further define an average loss function L : X → R as a result of the quantization. For some x ∈ X, where x belongs
to a quantization bin Bi whose representative state is yi (i.e. q(x) = yi ), the average loss function L(x) is defined as
178 8 Numerical and Approximation Methods
Z
L(x) := ∥x − x′ ∥ π̂y∗i (dx′ ) ∀x ∈ Bi , i = 1, · · · , M. (8.20)
Bi

That is, L(x) can be seen as the distance of the state x to the mean of the bin Bi under the measure π̂y∗i .
In the following, we present error analyses in finite state approximations defined in this section.

Finite State Approximations with Kernels Continuous in Total Variation under Expected Quantization Error
Bounds

In this section, we focus on an MDP model whose transition kernel is Lipschitz continuous in x (uniform in u) under
the total variation norm. This condition is somewhat different than the continuity of the transition kernel under the total
variation distance. Indeed, if we have a model as in Example 5.6-(ii), then we have the required Lipschitz continuity of
the transition kernel when f is Lipschitz continuous in x that is uniform in u and the density of the noise w is Lipschitz
continuous. The following assumptions are imposed on the system.

Assumption 8.2.4 (a) There exists a constant αc > 0 such that |c(x, u) − c(x′ , u)| ≤ αc ∥x − x′ ∥ for all x, x′ ∈ X and
for all u ∈ U.
(b) There exists a constant αT > 0 such that ∥T (·|x, u) − T (·|x′ , u)∥T V ≤ αT ∥x − x′ ∥ for all x, x′ ∈ X and for all
u ∈ U.

The first result gives an error bound for the approximate value function.

Theorem 8.2.5 [184, Theorem 4] Under Assumption 8.2.4, provided that c is bounded, we have for any initial state
x0 ∈ X
  ∞
ˆ ∗ βαT ∥c∥∞ X t
Jβ (x0 ) − Jβ (x0 ) ≤ αc + β sup ExT0,γ [L(Xt )] ,
1−β t=0 γ∈Γ

where L is defined in (8.20).

The following result provides an error bound for the approximate policy of the finite-state model when it is applied to
the original model.

Theorem 8.2.6 Under Assumption 8.2.4, we have for any initial state x0 ∈ X
 ∞
X
βαT ∥c∥∞
Jβ (x0 , γ̂) − Jβ∗ (x0 ) ≤ 2 αc + β t sup ExT0,γ [L(Xt )]
1−β t=0 γ∈Γ

where L is defined in (8.20) and γ̂ denotes the optimal policy of the finite-state approximate model given by (8.18)
extended to the state space X via the quantization function q.

Finite State Approximation with Kernels Continuous in Wasserstein Distance under Uniform Quantization
Error Bounds

In this section, we focus on the models with transition kernels that are Lipschitz continuous in x (uniform in u) under
the first order Wasserstein distance. If we have a model as in Example 5.6-(i), then we have the required Lipschitz
continuity of the transition kernel when f is Lipschitz continuous in x that is uniform in u. Here, instead of providing
an average loss bound using (8.20) as in Theorem 8.2.5, we will provide a uniform loss bound result, and we also
assume the state space to be compact. We first define
8.2 Approximation through Quantization of the State and the Action Spaces 179

L̄ := max sup ∥x − x′ ∥. (8.21)


i=1,...,M x,x′ ∈Bi

Here, L̄ is the largest diameter among the quantization bins.


The following is essentially Assumption 5.5.1, re-stated for the setup considered.

Assumption 8.2.5 (a) X is compact.


(b) There exists a constant αc > 0 such that |c(x, u) − c(x′ , u)| ≤ αc ∥x − x′ ∥ for all x, x′ ∈ X and for all u ∈ U.
(c) There exists a constant αT > 0 such that W1 (T (·|x, u), T (·|x′ , u)) ≤ αT ∥x − x′ ∥ for all x, x′ ∈ X and for all
u ∈ U.

Theorem 8.2.7 [184, Theorem 5] Let Assumption 8.2.5 hold and let X be compact. We have
αc
sup Jˆβ (x0 ) − Jβ∗ (x0 ) ≤ L̄
x0 ∈X (1 − βαT )(1 − β)

where L̄ is defined in (8.21).

The following result is similar to [270, Theorem 4.38] with a slightly improved bound.

Theorem 8.2.8 [184, Theorem 6] Let Assumption 8.2.5 hold and let X be compact. We have
2αc
sup Jβ (x0 , γ̂) − Jβ∗ (x0 ) ≤ L̄.
x0 ∈X (1 − β)2 (1 − βαT )

where L̄ is defined in (8.21) and γ̂ denotes the optimal policy of the finite-state approximate model extended to the state
space X via the quantization function q.

Finite State Approximation with Weakly Continuous Kernels and Asymptotic Convergence

In this section, we assume that X is σ-compact. That is, we can write X = ∪∞k=1 Bk where each Bk is compact. A finite
dimensional Euclidean space is an example of such a space. Additionally, in this section, we focus on the models with
transition kernels that are continuous only under the weak convergence topology. Here, instead of providing a rate of
convergence, we will provide an asymptotic result. Let the quantizer be such that the M th bin be the over-flow bin; that
is, the first M − 1 bins be the quantization of a compact set and the complement be assigned to BM . Let dX denote the
metric on X
To this end, let us define

L− := max sup dX (x, x′ ). (8.22)


i=1,...,M −1 x,x′ ∈Bi


Note that since X is σ-compact, for each M , one can find a partition {Bi }M i=1 of the state space X such that L → 0
SM −1 M −1
and i=1 Bi ↗ X as M → ∞. Note that BM = X \ (∪i=1 Bi ). In the following result, we assume that such a
sequence of partitions is used to obtain the finite-state approximate models.

Theorem 8.2.9 [270, Theorem 4.27] Under Assumption 8.2.1, we have for any compact K ⊂ X

sup Jˆβ (x0 ) − Jβ∗ (x0 ) → 0


x0 ∈K

and
180 8 Numerical and Approximation Methods

sup Jβ (x0 , γ̂) − Jβ∗ (x0 ) → 0


x0 ∈K

as L− → 0, where γ̂ denotes the optimal policy of the finite-state approximate model extended to the state space X via
the quantization function q.

We note that the result by [270, Theorem 4.27] is more general and applicable to unbounded cost functions as well.
Under the bounded cost in Assumption 8.2.1, [270, Theorem 4.27] implies Theorem 8.2.9 above.
One challenge to be addressed in the proofs of the results noted above is that in the quantized models (as an intermediate
step in the proof) we do not have the weak continuity condition for each of the quantized models: The issue is that
the value function in the dynamic programming update iterations is not continuous (and would only be piece-wise
continuous), and accordingly the finite models are not necessarily continuous in the action variables which violates the
measurable selection conditions noted in Section 5.2. Nonetheless, the machinery of universally measurable policies
(see Appendix C) can be utilized and the existence of optimal policies for the approximate kernels does not arise as
an immediate problem (see [275, p. 6-7]) for the proof of the theorem. Alternatively, one can first quantize the action
set and work on the approximate (finite-action) MDP, whose near optimality was established earlier. Note that for the
finite action setup, continuity of the kernels in actions always holds. Accordingly, we assume that the action set is finite
in the following analysis.

The Average Cost Case

Theorem 8.2.10 [270] [181, Theorem 4.2] [Average Cost] Suppose that Assumption 5.2.1 and Assumption 7.2.2 hold.
Then,

lim ∥J(·, γ̂) − J ∗ ∥ = 0. (8.23)


m→∞

where γ̂ denotes the optimal policy of the finite-state approximate model extended to the state space X via the quanti-
1
zation function q with diameter m .

One can also obtain explicit rates of convergence under further regularity, similar to the discounted cost setting studied
earlier.

Remark 8.4. We note that [270] imposes total variation continuity, but building on [182, Theorem 16] and adapting the
arguments in the proof of [181, Theorem 4.2] (see Section 12.4), the total variation condition continuity condition on
the kernel can be relaxed to weak Feller continuity, leading to the above. This will be studied in Chapter 12.

8.2.3 Finite Model MDP Approximation: Quantization of both the State and Action Spaces

We showed in Theorems 8.2.3 and 8.2.4 that any MDP with (infinite) compact action space and with bounded one-
stage cost function can be well approximated by an MDP with finite action space, for both the discounted cost and the
average cost cases.
Therefore, before discretizing the state space to compute near optimal policies, one can discretize, without loss of
generality, the action space U in advance on a finite grid using sufficiently large number of grid points. Then, near
optimality of finite models follow from the discussions in Section 8.2.2.

Remark 8.5 (An alternative direct argument via the convex analytic method for average cost criteria). For average
cost criteria, an alternative argument is presented in [16, Theorem 4.2], where it is shown that under either weak or
setwise continuity conditions, both finite state and action models are near optimal by considering the set of invariant
occupation measures under unique ergodicity conditions. Furthermore, under the conditions noted, deterministic and
stationary policies are shown to be dense among those that are randomized and stationary, in the sense that the cost
8.3 Numerical Methods for POMDPs 181

under any randomized stationary policy can be approximated arbitrarily well by deterministic and stationary policies.
Furthermore, the dense set of deterministic and stationary policies can be assumed to have finite range. A further utility
of this approach is that geometric ergodicity, e.g. via minorization, is not needed.

8.3 Numerical Methods for POMDPs

8.3.1 Near optimality of quantized policies under weak Feller or Wasserstein regularity of non-linear filters

As we have seen, any POMDP can be reduced to a completely observable Markov process [352], [262], whose states
are the posterior state distributions or ’beliefs‘ of the observer; that is, the state at time t is

πt ( · ) := P {Xt ∈ · |y0 , . . . , yt , u0 , . . . , ut−1 } ∈ P(X).

We called this conditional probability measure process the filter process. The filter process has state space P(X) and
action space U. Here, P(X) is equipped with the Borel σ-algebra generated by the topology of weak convergence. The
transition probability of the filter process is given in (6.27).
Accordingly, we have a fully observed belief-MDP. Now, by combining the approximation results in Section 8.2 and
reinforcement learning theoretic results to be presented in Section 9.3, together with the the weak Feller continuity
results presented in Section 6.3.2, we can conclude that the numerical methods can also be applied to POMDPs under
the conditions reported in Theorems 6.3.3 and 6.3.4 [127] [183].
This has explicitly been demonstrated in [276, Theorem 3], where also methods for quantizing probability measures
have been studied. This also applies under Wasserstein regularity of non-linear filter kernels, with explicit conditions
given in Theorem 6.3.5 [99]; see Theorem 8.2.7.
Accordingly, due to the weak Feller property of controlled non-linear filters, we can obtain rigorous approximation
results by quantizing probability measures.
In the following, we present an alternative approach.

8.3.2 Near-optimality of finite window policies under filter stability

One can also show that under filter stability (see Section 6.4), finite window policies are near optimal [187]. Consider
the following:
"∞ #
X
T ,γ
Jβ (µ, T , γ) = Eµ t
β c(xt , ut ) , Jβ∗ (µ, T ) = inf Jβ (µ, T , γ).
γ∈Γ
t=0

The question we ask is: suppose that instead of using all available history, we construct an approximate model using
the finite window information variables

ItN = {y[t−N,t] , u[t−N,t−1] }, if t ≥ N,


ItN = {y[0,t] , u[0,t−1] }, if 0 < t < N,
I0 = {y0 }, (8.24)

that is we observe the information variables through a window whose length is N . Suppose, we denote the optimal
value function of the approximate model by JβN and the approximate policy by γ N .
Inspired from filter stability, consider the following: For any time step t ≥ N and for a fixed observation realization
sequence y[0,t] and control action sequence u[0,t−1] , the state process can be viewed as
182 8 Numerical and Approximation Methods

P µ (xt ∈ ·|y[0,t] , u[0,t−1] ) = P πt−N − (Xt ∈ ·|y[t−N,t] , u[t−N,t−1] )

where

πt−N − (·) = P µ (xt−N ∈ ·|y[0,t−N −1] , u[0,t−N −1] ).

That is, we can view the state as the Bayesian update of πt−N− , the predictor at time t − N , using the observations
yt−N , . . . , yt . Notice that with this representation only the most recent N observation realizations are used for the
update and the past information of the observations is embedded in πt−N− . Consider the following set (state space):

Z = (π, y[0,N ] , u[0,N −1] ); π ∈ P(X), y[0,N ] ∈ YN +1 , u[0,N −1] ∈ UN




We place the product metric on this new space: weak convergence on the belief and usual metric on the measurements
and actions.
The approach is summarized in Figure 8.1.
Now, for a fixed π̂ ∈ P(X), consider the quantized state space:

Z N = (π̂, y[0,N ] , u[0,N −1] ); y[0,N ] ∈ YN +1 , u[0,N −1] ∈ UN




The idea is to quantize Z as follows: collapse all π to a fixed state π̂, define an approximate finite MDP and establish
performance bounds utilizing filter stability and the robustness approach presented earlier
In the following, we will assume that X is Rn for some n and that U, Y are finite sets.
The actual state space and the finite approximation are:

Z = (π, y[0,N ] , u[0,N −1] ); π ∈ P(X), y[0,N ] ∈ YN +1 , u[0,N −1] ∈ UN




Z N = (π̂, y[0,N ] , u[0,N −1] ); y[0,N ] ∈ YN +1 , u[0,N −1] ∈ UN




Define the map and F : Z → Z N , such that for (π, y[0,N ] , u[0,N −1] ) ∈ Z

F (π, y[0,N ] , u[0,N −1] ) = (π̂, y[0,N ] , u[0,N −1] ).

y0 , u0 ; y1 , u1 ; y2 , u2 ; y3 , u3 ; · · · ; yt−1 , ut−1 ; yt γ ut ADMISSIBLE POLICY

πt γ ut BELIEF REDUCTION

πt−N− ; yt−N , ut−N ; · · · , yt−1 , ut−1 ; yt γ ut FINITE WINDOW BELIEF REDUCTION

Quantizing the prior πt−N−

π̂; yt−N , ut−N ; · · · ; yt−1 , ut−1 ; yt γ ut APPROXIMATE FINITE WINDOW MDP

Fig. 8.1: Construction of the Finite-Window Approximate MDP from the Finite-Window Belief-MDP [188].
8.3 Numerical Methods for POMDPs 183

Using the map F and the finite set Z N , one can define a finite belief MDP, and construct a policy for this finite model,
by extending it, we can use the policy, say ϕ̃N for the original model.
The cost function for the approximate model is

ĉ(ẑtN , ut ) = ĉ(π̂, ItN , ut ) := c̃(ϕ(π̂, ItN ), ut )


Z
= c(xt , ut )P π̂ (dxt |yt , . . . , yt−N , ut−1 , . . . , ut−N ).
X

We define the controlled transition model for the approximate model by


 
η̂ N (ẑt+1
N
|ẑtN , ut ) = η̂ N (π̂, It+1
N
|π̂, ItN , ut ) := η̂ P(X), It+1
N
|π̂, ItN , ut . (8.25)

We will write Zπ̂N to make the dependence on π̂ and N more explicit.


N
For simplicity, if we assume N = 1, then the transitions can be rewritten for some It+1 = (ŷt+1 , ŷt , ût ) and ItN =
(yt , yt−1 , ut−1 )

η̂ N (π̂, ŷt+1 , ŷt , ût |π̂, yt , yt−1 , ut−1 , ut ) = η̂(P(X), ŷt+1 , ŷt , ût |π̂, yt , yt−1 , ut−1 , ut )
= 1{yt =ŷt ,ut =ût } P π̂ (ŷt+1 |yt , yt−1 , ut , ut−1 ). (8.26)

Denoting the optimal value function for the approximate model by JβN , we can write the following fixed point equation
 
X
JβN (ẑ N ) = min ĉ(ẑ N , u) + β JβN (ẑ1N )η̂ N (ẑ1N |ẑ N , u) . (8.27)
u∈U
ẑ1N ∈Ẑπ̂
N

E.g., for N = 1, we rewrite the fixed point equation for some ẑ0N = (π̂, y1 , y0 , u0 ) as
 
X
JβN (π̂, y1 , y0 , u0 ) = min ĉ(π̂, y1 , y0 , u0 , u1 ) + β JβN (π̂, y2 , y1 , u1 )P π̂ (y2 |y1 , y0 , u1 , u0 ) . (8.28)
u1 ∈U
y2 ∈Y

We can now investigate the following approximation error terms:

|J˜βN (ẑ) − Jβ∗ (ẑ)|, Jβ (ẑ, ϕ̃N ) − Jβ∗ (ẑ).

The first one is the difference between the optimal value function of the original model and that for the approximate
model. The second term is the performance loss due to the policy calculated for the approximate model being applied
to the true model.
Using the approaches presented in Section 8.2 and what is to be presented in Chapter 12 (building on [182, 186], we
can show that the loss is related to the term:

Lt :=
h −
i
sup Eπγ̂− ∥P πt (Xt+N ∈ ·|Y[t,t+N ] , U[t,t+N −1] ) − P π̂ (Xt+N ∈ ·|Y[t,t+N ] , U[t,t+N −1] )∥T V (8.29)
0
γ̂∈Γ̂

Notice that this term is directly related to filter stability studied in Chapter 6.

Theorem 8.3.1 [187, 188] [Continuity of Value Functions] For ẑ0 = (π0− , I0N ), if a policy γ̂ acts on the first N step of
the process which produces I0N , we then have
184 8 Numerical and Approximation Methods

h i ∥c∥∞ X t
Eπγ̂− ˜N ∗ N
Jβ (ẑ0 ) − Jβ (ẑ0 ) |I0 ≤ β Lt
0 (1 − β) t=0

Theorem 8.3.2 [187, 188] [Robustness of Approximate Finite Window Model Solution applied to Actual Model] For
ẑ0 = (π0− , I0N ), with a policy γ̂ acting on the first N steps

h i 2∥c∥∞ X t
Eπγ̂− Jβ (ẑ0 , ϕ̃N ) − Jβ∗ (ẑ0 ) |I0N ≤ β Lt .
0 (1 − β) t=0

As one example, we now show that the term Lt can be bounded by the filter stability result presented in Theorem 6.4.1.
Recall that this states that

E µ,γ [∥πnµ,γ − πnν,γ ∥T V ] ≤ 2αn .

which holds uniformly for all µ ≪ ν where α := (1 − δ̃(T ))(2 − δ(Q)).


Since δ̃(T ) is a uniform Dobrushin coefficient over all control actions, the above bound is valid under any control
policy. Thus we have that
h −
i
Lt = sup Eπγ̂− ∥P πt (Xt+N ∈ ·|Y[t,t+N ] , U[t,t+N −1] ) − P π̂ (Xt+N ∈ ·|Y[t,t+N ] , U[t,t+N −1] )∥T V
γ∈Γ 0

N
≤ 2α (8.30)

Theorem 8.3.3 [187, 188] Assume the following holds:


(i) The exponential filter stability condition applies:

α := (1 − δ̃(T ))(2 − δ(Q)) < 1 (8.31)

(ii) The transition kernel T is dominated, i.e. there exists a dominating measure π̂ ∈ P(X) such that for every x ∈ X
and u ∈ U, T (·|x, u) ≪ π̂(·).
Then, by choosing the dominating measure π̂ for the approximate model,
h i 4∥c∥∞ N
Eπγ̂− Jβ (ẑ0 , ϕ̃N ) − Jβ∗ (ẑ0 ) |I0N ≤ α . (8.32)
0 (1 − β)2

Instead of the Dobrushin based analysis, the more relaxed stability condition presented in Example 6.9 can also be
adopted, though not leading to an exponential convergence rate in the memory size. In that case, Lt → 0 asymptotically
and thus (8.32) converges to 0 as N increases only asymptotically, without a geometric convergence rate in N .
Via a somewhat different, and more direct, derivation, [187, Section 4.2 and Theorem 17] presented the following
alternative condition involving sample path-wise uniform filter stability term

L̄N
T V := sup sup P z (·|y[0,N ] , u[0,N −1] ) − P z (·|y[0,N ] , u[0,N −1] ) , (8.33)
z∈P(X) y[0,N ] ,u[0,N −1] TV

to show the following uniform error bound:

2(1 + (αZ − 1)β)


sup Jβ (z, γN ) − Jβ∗ (z) ≤ ∥c∥∞ L̄N
TV (8.34)
z (1 − β)3 (1 − αZ β)

for all β ∈ (0, 1) under a contraction condition, for some constant αZ defined in [187]. Additionally, [187, Theorem
9] provided conditions where the error is in the bounded-Lipschitz metric (which is equivalent to the Wasserstein-1
metric when the state space X is compact), however these were only applicable for a restrictive subset of the discount
8.4 Bibliographic Notes 185

parameter β. On the other hand, the bound in (8.32) is in expectation whereas the bound in (8.34) is uniform, and thus
the results are complementary.
As a complementary condition, via the Birkhoff-Hopf theorem, a controlled version of a contraction via the Hilbert
metric [145] can be utilized [100]: Recall that

F (z, y, u)(·) = Pr {Xk+1 ∈ · | Zk = z, Yk+1 = y, Uk = u}

Assumption 8.3.1 1. Q(y|x) ≥ ϵ > 0 for every x ∈ X and y ∈ Y.


2. The transition kernel T (.|., u) is a mixing kernel (see Definition 6.4.8) for every u ∈ U.

Lemma 8.3.1 [1] Under Assumption 8.3.1, there exists a constant r < 1 such that

h(F (µ, y, u), F (ν, y, u)) ≤ rh(µ, ν) (8.35)

1−ϵ2u ϵ
for every comparable µ, ν ∈ P(X) and for every u ∈ U and y ∈ Y. Here r = 1+ϵ2u ϵ , ϵu is the mixing constant of the
kernel T (.|., u).

Theorem 8.3.4 [100] Under Assumption 8.3.1, there exists a constant r < 1 and K such that
N −1
L̄N
TV ≤ r K. (8.36)

2 1−ϵ2u ϵ
Here, K = log 3 sup h(Z1 , Z1∗ ) and r = supu∈U 1+ϵ2u ϵ .

Corollary 8.6. [100] Under Assumption 8.3.1, there exists a constant r < 1 and K such that

− − 2∥c∥∞ N −1
, T , γ N ) − Jβ∗ (πN , T )|I0N ≤
 
E Jβ (πN r K. (8.37)
(1 − β)2

2 1−ϵ2u ϵ
Here, K = log 3 sup h(Z1 , Z1∗ ) and r = supu∈U 1+ϵ2u ϵ .

Despite these consequential approximation results, implementing the above is still tedious; though possible. Can rein-
forcement learning be feasible? Can we view the finite history as an effective state? We will address this question in
the following chapter.

8.4 Bibliographic Notes

For computational and learning methods, there is an extensive literature, where various approaches have been devel-
oped. A partial list of these techniques is as follows: approximate dynamic programming, approximate value or policy
iteration, simulation-based techniques, neuro-dynamic programming (or reinforcement learning), state aggregation,
etc. [39, 41, 87, 111].
We refer the reader to the monograph [270] for a general treatment of approximation results along what has been
presented for continuous spaces; these also build on on [184, 269, 272, 273, 275]. presentation, only focus on the setup
where the state spaces considered are compact. A generalization of some of the approximation results are presented
in [186] in view of robustness properties.
186 8 Numerical and Approximation Methods

8.5 Exercises

Exercise 8.5.1 Consider a controlled Markov chain with state space X = {0, 1}, action space U = {0, 1}, and
transition kernel for t ∈ Z+ :

P (xt+1 = 1|xt = 0, ut = 1) = P (xt+1 = 1|xt = 1, ut = 1) = α

P (xt+1 = 1|xt = 0, ut = 0) = P (xt+1 = 1|xt = 1, ut = 0) = 1 − α.


where α ∈ (0, 1). Let a cost function c(x, u), with c : X × U → R+ be given by

c(0, 1) = c(0, 0) = 1 c(1, 0) = c(1, 1) = 2.

Suppose that the goal is to minimize the quantity



X
E0γ [ β t c(xt , ut )],
t=0

for a fixed β ∈ (0, 1), over all admissible policies γ ∈ ΓA .


Find an optimal policy and the optimal expected cost explicitly, as a function of α, β (note that the initial condition is
x0 = 0).

Exercise 8.5.2 Consider the following problem: Let X = {1, 2}, U = {1, 2}, where X denotes whether a fading
channel is in a good state (x = 2) or a bad state (x = 1). There exists an encoder who can either try to use the channel
(u = 2) or not use the channel (u = 1). The goal of the encoder is send information across the channel.
Suppose that the encoder’s cost (to be minimized) is given by:

c(x, u) = −1{x=2,u=2} + α(u − 1),

for α = 1/2 (if you view this as a maximization problem, you can see that the goal is to maximize information trans-
mission efficiency subject to a cost involving an attempt to use the channel; the model can be made more complicated
but the idea is that when the channel state is good, u = 2 can represent a channel input which contains data to be
transmitted and u = 1 denotes that the channel is not used).
Suppose that the transition kernel is given by:

P (xt+1 = 2|xt = 2, ut = 2) = 0.8, P (xt+1 = 1|xt = 2, ut = 2) = 0.2

P (xt+1 = 2|xt = 2, ut = 1) = 0.2, P (xt+1 = 1|xt = 2, ut = 1) = 0.8


P (xt+1 = 2|xt = 1, ut = 2) = 0.5, P (xt+1 = 1|xt = 1, ut = 2) = 0.5
P (xt+1 = 2|xt = 1, ut = 1) = 0.9, P (xt+1 = 1|xt = 1, ut = 1) = 0.1

We will consider either a discounted cost criterion for some β ∈ (0, 1) (you can fix an arbitrary value)

X∞
inf Exγ [ β t c(xt , ut )] (8.38)
γ
t=0

or the average cost criterion


T −1
1 γ X
inf lim sup Ex [ c(xt , ut )]. (8.39)
γ T →∞ T t=0

a) Using Matlab or some other program, obtain a solution to the problem given above in (9.34) through the following:
8.5 Exercises 187

(i) Policy Iteration


(ii) Value Iteration.
b) Consider the criterion given in (9.35). Apply the convex analytic method, by solving the corresponding linear pro-
gram, to find the optimal policy. In Matlab, the command linprog can be used to solve linear programming problems.
See (7.42).

Exercise 8.5.3 (Convex Analytic Method for Discounted Cost) Consider the convex analytic method for a discounted
cost problem, where the expected occupation measures are defined, for a given γ ∈ ΓA as

X
η(A × B) = β k P γ (Xk ∈ A, Uk ∈ B)
k=0

Show that the set of all such expected occupation measures η are equivalent to the set of all expected occupation
measures achieved by stationary (and randomized) γ, that is, by γ ∈ ΓSR .
Hint: Obtain a recursive equation involving

X
η(A × U) = P (X0 ∈ A) + β β t−1 P γ (Xk ∈ A)
k=1

X Z
= P (X0 ∈ A) + β β k−1 T (A|x, u)P γ (Xk−1 ∈ dx, Uk−1 ∈ du)
k=1
Z Z ∞
X
= P (X0 ∈ A) + β T (A|x, u) β k−1 P γ (Xk−1 ∈ dx, Uk−1 ∈ du)
k=1
Z Z
= P (X0 ∈ A) + β T (A|x, u)η(dx, du)

Then, note the similarity with (8.7), in that this equation is also satisfied by selecting a stationary control κ defined
almost everywhere with κ(du|x) = dη(dx,du)
dη(dx) (x).

Exercise 8.5.4 Let c : X × U → R+ be bounded, where X is the state space and U is the action space for a controlled
stochastic system. Suppose that under a stationary policy γ, the expected discounted cost, for β < 1, is given by

X Z
Jβ (x, γ) := Exγ [ β k c(xk , γ(xk ))] = c(x, γ(x)) + β Jβ (xt+1 , γ)T (dxt+1 |xt = x, ut = γ(x))
k=0

Let f1 and f2 be two stationary policies. Define a third policy, g, as:

g(x) = f1 (x)1{x∈C} + f2 (x)1{x∈X\C}

where
C = {x : Jβ (x, f1 ) ≤ Jβ (x, f2 )}
and X \ C denotes the complement of this set.
Show that Jβ (x, g) ≤ Jβ (x, f1 ) and Jβ (x, g) ≤ Jβ (x, f2 ) for all x ∈ X.

Exercise 8.5.5 Consider a controlled Markov chain with state space X = {0, 1}, action space U = {0, 1}, and
transition kernel for t ∈ Z+ :
P (xt+1 = 1|xt = 0, ut = 1) = 1
1
P (xt+1 = 1|xt = 0, ut = 0) =
2
188 8 Numerical and Approximation Methods

1
P (xt+1 = 1|xt = 1, ut = 0) = P (xt+1 = 1|xt = 1, ut = 1) = .
2
Let a cost function c : X × U → R+ be given by
1
c(0, 1) = , c(0, 0) = 1
2
5
c(1, 0) = , c(1, 1) = 2
4

Suppose that the goal is to minimize the quantity



X 1
E0γ [ ( )t c(xt , ut )],
t=0
2

over all admissible policies γ ∈ ΓA .


Find an optimal policy using Policy Iteration.
Note. Please note that you can only use a pen and paper for this problem. Note that for an invertible 2x2 matrix
 
ab
A=
cd

we have  
d −b
−c a
A−1 =
ad − bc
9

Reinforcement Learning

In this chapter, we will study stochastic learning and reinforcement learning methods, first in the context of finite space
MDPs and later on for general continuous (standard Borel) space MDPs and POMDPs.

9.1 Stochastic Learning Algorithms and the Q-Learning Algorithm

In some Markov Decision Problems (MDPs), one does not know the true transition kernel or the cost function, and
may wish to use past data to obtain an asymptotically optimal solution (that is, via learning from past data). In some
problems, this may be used as an efficient numerical method to obtain approximately optimal solutions. There may
also be setups where a prior probabilistic knowledge on the system dynamics may be used to learn the true system. In
particular, one may apply Bayesian (probabilistically driven given some prior information) or non-Bayesian (primarily
empirical, without assuming a prior probabilistic model) methods.
An important class of non-Bayesian methods are known as stochastic approximation algorithms: such approximation
methods are used extensively in many application areas. A typical stochastic approximation algorithm has the following
form

xt+1 = xt + αt (F (xt ) − xt + wt ) (9.1)

where wt is a zero-mean noise variable, xt is a stochastic process and wt is some driving noise. The goal is to arrive at
a point x∗ which satisfies x∗ = F (x∗ ), where F may correspond to an optimality condition.

Exercise 9.1.1 (Stochastic gradient descent) We revisit the stochastic gradient descent algorithm discussed in Theo-
rem 4.3.4 noting the similarity with (9.1). Consider a convex function f : Rn → R and denote the set of minima of f
by X ∗ . Let a sequence of iterates be given by
 
xk+1 = xk − γk ∇x f (xk ) + nk , x0 ∈ Rn , (9.2)

where the random variables nkP


are zero-mean, orthogonal to one another, and have their second moments uniformly
bounded. If k γk = ∞ and k γk2 < ∞, the sequence of iterates (9.2) converges almost surely to some element
P
x∗ ∈ X ∗

9.1.1 Q-Learning

Q−learning [26, 41, 300, 303, 306, 323] is a stochastic approximation algorithm used for fully observed finite space
MDPs that does not require the knowledge of the transition kernel, or even the cost (or reward) function for its imple-
mentation. In this algorithm, the incurred per-stage cost variable is observed through simulation of a single sample path.
190 9 Reinforcement Learning

In this context, one may reflect on the way humans respond to experience, for example a baby learning to experiment
with gravity without knowing physics and therefore physical models at all!
When the state and action spaces are finite, under mild conditions regarding infinitely often hits for all state-action
pairs, this algorithm is known to converge to the optimal cost. We now discuss this algorithm.
Consider a Markov Decision Problem with finite state and action sets with the criterion given in (8.1) for some β ∈
(0, 1).
Let Q : X × U → R denote the Q-factor of the controller. Let us assume that the decision maker applies an arbitrary
admissible policy γ and updates its Q-factors as follows for t ≥ 0,
 
Qt+1 (x, u) = Qt (x, u) + αt (x, u) c(x, u) + β min Qt (Xt+1 , v) − Qt (x, u) (9.3)
v

where the initial condition Q0 is given, αt (x, u) is the step-size for (x, u) at time t, ut is chosen arbitrarily -e.g.,
according to some random exploration policy γ- as long as some technical conditions noted below hold, and the
random state Xt+1 ∼ P (Xt+1 ∈ · |Xt = x, Ut = u). It is assumed that, for all (x, u), t ≥ 0, the following hold

Assumption 9.1.1 For all (x, u), t ≥ 0,


(i) αt (x, u) ∈ [0, 1]
(ii) αt (x, u) = 0 unless (x, u) = (xt , ut )
(iii)αt (x, u) is a (deterministic) function of (x0 , u0 ), . . . , (xt , ut ).
P
(iv) t≥0 αt (x, u) = ∞, almost surely

(v) t≥0 αt2 (x, u) ≤ C, almost surely, for some (deterministic) constant C < ∞.
P

A common way to select α coefficients in the algorithm is to take for every (x, u) pair:
1
αt (x, u) = Pt
1+ 1
k=0 {Xk =x,Uk =u}

The selection of the control actions for each state can be arbitrary, as long as the assumptions above are guaranteed to
hold.
Let F be an operator acting on the Q factors defined by:
X
F (Q)(x, u) = c(x, u) + β T (x′ |x, u) min Q(x′ , v), (9.4)
v
x′

where, as before, T (x′ |x, u) = P (x1 = x′ |x0 = x, u0 = u) is the transition kernel. Consider the following fixed point
equation.
X
Q∗ (x, u) = F (Q∗ )(x, u) = c(x, u) + β T (x′ |x, u) min Q∗ (x′ , v) (9.5)
v
x′

whose existence and uniqueness follow essentially identically from arguments used in the contraction analysis utilized
in Chapter 5 (see Theorem 5.5.2), by using the norm ∥Q∥∞ = max(x,u) |Q(x, u)|: Since for every realization xt+1 of
Xt+1 , | minv Qt (xt+1 , v) − minv′ Q∗t (xt+1 , v ′ )| ≤ maxv |Qt (xt+1 , v) − Q∗t (xt+1 , v)|, we have that

|F (Qt )(x, u) − F (Q∗ )(x, u)| ≤ β∥Qt − Q∗ ∥∞ := β max |Qt (x, u) − Q∗ (x, u)| (9.6)
x,u

Now, note that we can write (9.3) as


9.1 Stochastic Learning Algorithms and the Q-Learning Algorithm 191

Qt+1 (x, u) = Qt (x, u) + αt (x, u) F (Qt )(x, u) − Qt (x, u)
 
+ c(x, u) + β min Qt (xt+1 , v) − F (Qt )(x, u)
v

(9.7)

which is in the same form as (9.1) since


 
wt := c(xt , ut ) + β min Qt (xt+1 , v) − F (Qt )(xt , ut ) , (9.8)
v

conditioned on the filtration generated by {xt , ut } up to time t, is a zero-mean random variable.


Let us write (9.7) as
  
Qt+1 (x, u) = (1 − αt (x, u))Qt (x, u) + αt (x, u) F (Qt )(x, u) + c(x, u) + β min Qt (xt+1 , v) − F (Qt )(x, u)
v

and by (9.5)

Qt+1 (x, u) − Q∗ (x, u)



= (1 − αt (x, u))(Qt (x, u) − Q∗ (x, u)) + αt (x, u) F (Qt )(x, u) − F (Q∗ )(x, u)
 
+ c(x, u) + β min Qt (xt+1 , v) − F (Qt )(x, u) (9.9)
v

Theorem 9.1.1 (i) Under Assumption 9.1.1, the algorithm (9.3) converges almost surely to Q∗ .
(ii) A stationary policy f ∗ which satisfies minu Q∗ (x, u) = Q∗ (x, f ∗ (x)) is an optimal policy.

Proof. (i) From (9.9), the process Qt satisfies the following form, with St = Qt − Q∗t :
 
St+1 (x, u) = (1 − αt (x, u))St (x, u) + αt (x, u) (F (Qt )(x, u) − F (Q∗ )(x, u)) + wt ,

where {αt } satisfies Assumption 9.1.1 and wt is given in (9.8).


We will consider the following two parallel dynamics, as in [177, Theorem 1]:
a
St+1 (x, u) = (1 − αt (x, u))Sta (x, u) + αt (x, u)wt , (9.10)
 
b b ∗
St+1 (x, u) = (1 − αt (x, u))St (x, u) + αt (x, u) F (Qt )(x, u) − F (Q )(x, u) . (9.11)

a b
We have St (x, u) = St+1 (x, u) + St+1 (x, u). We will study each of these two additive terms separately.
a
The first step is to show that St+1 (x, u) → 0 almost surely. We will show this further below.
Assume then for now that Sta (x, u) → 0 almost surely. We then focus on St+1
b a
(x, u) + St+1 (x, u), using the fact that,
by (9.6)
∥(F (Qt )(·, ·) − F (Q∗ )(·, ·))∥∞ ≤ β∥St ∥∞ ≤ β∥St+1
a
∥∞ + β∥St+1b
∥∞
a
Almost surely for sample paths ω, for every ϵ > 0, there exists N (ω) such that for t ≥ N (ω), ∥St+1 (ω)∥∞ ≤ ϵ (where,
in the following, we suppress the sample path dependence and thus omit ω). In the following, we assume that t ≥ N .
1
Now, for some M large enough, let β̂ := β(1 + M ) < 1 and for ∥Stb ∥∞ > M ϵ note that

β∥Stb (x, u) + ϵ∥ ≤ β̂∥Stb ∥∞


192 9 Reinforcement Learning

and
 
b
|St+1 (x, u)| ≤ (1 − αt (x, u))|Stb (x, u)| + αt (x, u) F (Qt )(x, u) − F (Q∗ )(x, u)

≤ (1 − αt (x, u))|Stb (x, u)| + αt (x, u)(β∥St+1


a b
∥ + β∥St+1 ∥) (9.12)
≤ (1 − αt (x, u))|Stb (x, u)| + αt (x, u)β̂∥Stb ∥∞ (9.13)
< ∥Stb ∥∞ (9.14)

Hence maxx,u (Stb (x, u)) monotonically decreases for ∥Stb ∥∞ > M ϵ leading to two possibilities: it either gets below
M ϵ or it never gets below M ϵ in which case by the monotone non-decreasing property it will converge to some number,
say M1 with M1 ≥ M ϵ.
Now, if the former is the case: once ∥Stb ∥∞ ≤ M ϵ we can show via (9.12) and β(M + 1)/M < 1 that it will remain
there thereafter.
We now show that the latter, that is with the limit being M1 ≥ M ϵ, is not possible. The relation
b
|St+1 (x, u)| ≤ (1 − αt (x, u))|Stb (x, u)| + αt (x, u)β̂∥Stb ∥∞

implies that (via an argument similar to what is known as Grönwall’s lemma, as can be inductively shown) the solution
is bounded from above by the solution to the equation
b
|St+1 (x, u)| = (1 − αt (x, u))|Stb (x, u)| + αt (x, u)β̂∥Stb ∥∞

which can be shown to converge to zero. This follows from the reasoning that, for any fixed D, the iterate
b
|St+1 (x, u)| = (1 − αt (x, u))|Stb (x, u)| + αt (x, u)β̂D

converges to β̂D; this follows since the effects of the initial condition diminish by the summability of αt (see Exercise
3.5.7) and hence there exists only one limit solution, which by inspection will be equal to β̂D (the uniqueness of the
solution can be shown by subtracting β̂D from the iterates, whose limit would be zero for any initialization). Therefore,
if there is an upper bound D0 on the iterates, the bounds for future iterations eventually get smaller and smaller than
β̂D0 +δ =: D1 for any arbitrarily small δ > 0, and as time progresses by an inductive reasoning, eventually the iterates
would have to converge to zero (see also [177, Proof of Lemma 3] or [306, p. 196]).
Thus, for any ϵ > 0, for large enough t, we have that ∥Stb ∥∞ ≤ M ϵ. Since ϵ > 0 is arbitrary, the convergence result
follows.
We now discuss Sta 1 .

1
We note that an alternative, more direct, argument, due [188, Theorem 4.1], is also possible under additional structure on the random
exploration policy used to generate the actions and if the learning rate leads to averaging dynamics as in Exercise 9.1.2; see [188, Theorem
4.1] for the analysis to follow and [190, Theorem 2.1]: In particular, if the policy adopted to randomly generate actions leads to a positive
Harris recurrent Markov chain, then the following more direct argument is possible to establish convergence without apriori showing
boundedness of the iterates: Write
  
Qt+1 (x, u) = (1 − αt (x, u))Qt (x, u) + αt (x, u) F (Qt )(x, u) + c(x, u) + β min Qt (xt+1 , v) − F (Qt )(x, u)
v

and

Qt+1 (x, u) − Q∗ (x, u) = (1 − αt (x, u))(Qt (x, u) − Q∗ (x, u))


 
+αt (x, u) F (Qt )(x, u) − F (Q∗ )(x, u) + β min Qt (xt+1 , v) − βE[min Qt (xt+1 , v)|xt = x, ut = u]
v v

(9.15)

With
9.1 Stochastic Learning Algorithms and the Q-Learning Algorithm 193

Taking the square of Sta , we obtain:


a
E[(St+1 (x, u))2 |Ft ] ≤ (Sta (x, u))2 − 2αt (Sta (x, u))2 + αt2 (Sta (x, u))2 + αt2 (x, u)wt2 (9.19)

First, by an argument identical to that used in the proof of the Comparison Theorem (Theorem 4.2.3), we have that for
any T > 0:
T
X −1 T
X −1
E[ (2αt − αt2 )(Sta (x, u))2 ≤ (S0a (x, u))2 + E[ αt2 (x, u)wt2 ]
t=0 t=0
≤ (S0a (x, u))2 + C(sup wt2 ), (9.20)
t

where we use Assumption 9.1.1(v). We now show that (supt wt2 ) is uniformly bounded: By (9.7), we have that

Qt+1 (x, u) = (1 − αt (x, u))Qt (x, u) + αt (x, u)(c(x, u) + β min Qt (xt+1 , v))
v

which implies that


|Qt+1 (x, u)| ≤ (1 − αt (x, u))∥Qt ∥∞ + αt (x, u)(c(x, u) + β∥Qt ∥∞ )
∥c∥∞
Now, if ∥Qt ∥∞ > 1−β =: L1 , we have that

|Qt+1 (x, u)| ≤ (1−αt (x, u))∥Qt ∥∞ +αt (x, u)(c(x, u)+β∥Qt ∥∞ ) = ∥Qt ∥∞ +αt (x, u)(c(x, u)+(β−1)∥Qt ∥∞ ) < ∥Qt ∥∞ ,

And thus, since this holds for all (x, u) pairs, ∥Qt+1 ∥∞ < ∥Qt ∥∞ whenever ∥Qt ∥∞ > L1 and hence |Qt ∥∞ would
be a decreasing sequence as long as it is above L1 . On the other hand, if ∥Qt ∥∞ ≤ L1 , then ∥Qt+1 ∥∞ ≤ L1 as well.
These imply that ∥Qt ∥∞ is uniformly bounded almost surely. As a consequence wt is also bounded, uniformly over
time. Thus, the right hand side of (9.20) is bounded.
Furthermore, by re-writing (9.19), in the expression
a
E[(St+1 (x, u))2 |Ft ] ≤ (Sta (x, u))2 − (2αt − αt2 )(Sta (x, u))2 + αt2 (x, u)wt2 ,

rt (x, u) := β min Qt (xt+1 , v) − βE[min Qt (xt+1 , v)|xt = x, ut = u].


v v

rt∗ (x, u) := β min Q (xt+1 , v) − βE[min Q∗ (xt+1 , v)|xt = x, ut = u]



v v

Qt+1 (x, u) − Q∗ (x, u) = (1 − αt (x, u))(Qt (x, u) − Q∗ (x, u))


 
+αt (x, u) F (Qt )(x, u) − F (Q∗ )(x, u) + rt∗ (x, u) + (rt (x, u) − rt∗ (x, u))

Obtain the sum:


a
Rt+1 (x, u) = (1 − αt (x, u))Sta (x, u) + αt (x, u)r∗ (x, u)t , (9.16)
 
b
Rt+1 (x, u) = (1 − αt (x, u))Stb (x, u) + αt (x, u) (rt (x, u) − rt∗ (x, u) , (9.17)
 
c
Rt+1 (x, u) = (1 − αt (x, u))Stb (x, u) + αt (x, u) F (Qt )(x, u) − F (Q∗ )(x, u) , (9.18)

a b c
St (x, u) = Rt+1 (x, u) + Rt+1 (x, u) + Rt+1 (x, u). By ergodicity, Rta → 0 almost surely by averaging and the positive Harris
recurrence under the exploration policy (by the additional assumption noted). Note that

|(rt (x, u) − rt∗ (x, u) + F (Qt )(x, u) − F (Q∗ )(x, u)| = |β min Qt (xt+1 , v) − β min Q∗ (xt+1 , v)| ≤ β∥Qt − Q∗t ∥∞
v v

and thus,
b
Rt+1 c
(x, u) + Rt+1 (x, u) ≤ β∥Qt − Q∗t ∥∞
b b c
As a result, by replacing St+1 with Rt+1 + Rt+1 we can trace the proof steps presented above without the analysis on Sta to follow.
194 9 Reinforcement Learning

the term αt2 (y, u)wt2 is finite almost surely. This implies, by Theorem 4.3.1, that Sta converges to some random variable
almost surely. The inequality (9.20) then implies that this limit must be zero: Suppose not; since αt is not summable,
there exists an infinite sequence of times so that each summation of αt between the times is bounded from below by a
positive constant. Through this, if (Sta )2 were not to converge to zero (given that it does converge toPsomething else), it
would remain above a positive constant after a sufficiently large time, and then it would follow that t (2αt − αt2 )Sta 2
would not remain bounded. Therefore, if this were to happen with non-zero measure, the expectation of this term would
be unbounded, which in turn would, as T → ∞, violate (9.20). You can also, alternatively, build on Theorem 4.3.2.
(ii) Now, consider X
Q∗ (x, u) = F (Q∗ )(x, u) = c(x, u) + β P (x′ |x, u) min Q∗ (x′ , v)
v
x′

Note that the minimum of u, for each x, is essentially the solution to the Discounted Cost Optimality Equation studied
in Chapter 5 (see Theorem 5.5.2). Hence, the stationary policy {f ∗ } is optimal. ⋄
We also refer the reader to the proof of [177, Theorem 1] for an alternative proof.

Exercise 9.1.2 To gain some further intuition, and also a more direct proof for the case where the αk term is taken
as k1 (or more precisely; for every (x, u) pair: αt (x, u) = Pt 1
), consider the following averaging
1+ 1{xk =x,uk =u}
k=0
dynamics: Let at be a sequence of scalars and define:
T −1
1 X
sT = ak
T
k=0

Observe that for T > 1, T sT = (T − 1)sT −1 + aT −1 which leads to


1
sT = sT −1 + (aT −1 − sT −1 )
T
In view of this observation, conclude that with αk in Assumption 9.1.1 taken as k1 , we have an averaging dynamics.
One may interpret the Q-learning algorithm and its convergence properties with this insight. Furthermore, one can see
that the convergence holds under more relaxed conditions, this will be utilized in Sections 9.2.2 and 9.3.

Remark 9.1. Via studying (9.10) and (9.13) separately, one can also arrive at convergence rates as a function of the
number of iterates, see [5, 120, 299].

9.1.2 Reinforcement Learning for the Average Cost Criterion

For the average cost criterion, we can follow two approaches: (i) One is to utilize near optimality of discounted criterion
policies as shown in Theorem 7.3.5 and Theorem 7.3.6; and (ii) another approach is via directly applying an algorithm
tailored for the average cost criterion. We note here that the average cost setup is typically more challenging due to the
lack of contraction properties. Nonetheless, as we have seen earlier in Chapter 7, one can follow contraction updates
for the average cost criterion as well. See [189] for algorithms and an extensive review.
We also note that the references [2, 149] are among the earliest studies that provide convergent learning algorithms
based on relative value iteration, and the convergence of these algorithms in these studies have been established via the
ODE method [64] for finite models.

9.1.3 Synchronous Q-Learning

The algorithm above is known as the asynchronous Q-learning algorithm: At any given time only a single (x, u) pair
can be updated. In some applications, where extensive data or simulations are available, one can simultaneously update
multiple or all state-action pairs, leading to a synchronous version. The analysis above is applicable to the synchronous
9.2 Reinforcement Learning Methods for POMDPs 195

setup, but the synchronous setup often allows for a more direct stochastic analysis especially for the average cost
criterion.

9.2 Reinforcement Learning Methods for POMDPs

For the analysis in this section, please first recall the discussion in Section 8.3.1.

9.2.1 Near optimality of quantized policies under weak Feller property of non-linear filters

As we have seen, any POMDP can be reduced to a completely observable Markov process, whose states are the
posterior state distributions or ’beliefs‘ of the observer; that is, the state at time t is

πt ( · ) := P {Xt ∈ · |y0 , . . . , yt , u0 , . . . , ut−1 } ∈ P(X).

As discussed earlier, this conditional probability measure process is the filter process. The filter process has state
space P(X) and action space U. Here, P(X) is equipped with the Borel σ-algebra generated by the topology of weak
convergence. The transition probability of the filter process is given in (6.27).
Accordingly, we have a fully observed belief-MDP. Now, by combining the approximation results in Section 8.2 and
reinforcement learning theoretic results to be presented in Section 9.3, together with the the weak Feller continuity
results presented in Section 6.3.2, we can conclude that the numerical methods can also be applied to POMDPs under
the conditions reported in Theorems 6.3.3 and 6.3.4 [127] [183].
Accordingly, due to the weak Feller property of controlled non-linear filters, we can apply quantized Q-learning, to be
introduced in Section 9.3, to also belief-based models to also arrive at near optimality of control policies. However, one
should note that some subtleties with regard to unique ergodicity properties arise; see [190].

9.2.2 Near-optimality of finite window policies under filter stability and Q-learning convergence

As an alternative approach, we also saw earlier in Section 8.3.2 that finite memory policies are near optimal under filter
stability conditions; see the program in Figure 8.1. We now study the reinforcement learning implementation of this
approach.
Recall that under mild conditions, for finite state and action MDPs, the Q-learning algorithm given in (9.3) converges
to a fixed point which leads to the Discounted Cost Optimality Equation (DCOE).
Learning in POMDPs is challenging, mainly due to the non-Markovian behavior of the observation process. For
POMDPs, an attempt may be to study the iterations given by

Qk+1 (yk , uk ) = (1 − αk (yk , uk ))Qk (yk , uk )


 
+ αk (yk , uk ) Ck (yk , uk ) + β min Qk (Yk+1 , v)
v

However, the observation process yt is not a controlled Markov process and the cost that is realized is c(xk , uk ), which
is not a function of yk and uk only. A two-part question then is the following:
(i) Would the Q-learning iterates for such a setup indeed converge?,
(ii) And, if they do converge, where do they converge to?
The answer to the first part of the question is positive under mild conditions [290] and [302]; and the answer to the
second part of the question is that under filter stability conditions, the convergence is to near optimality with an explicit
error bound between the performance loss and the memory window size.
196 9 Reinforcement Learning

Thus, to answer this, we consider a generalization using a finite window and use again filter stability [188] :
Assume that we start keeping track of the last N + 1 observations and the last N control action variables after at least
N + 1 time steps. That is, at time t, we keep track of the information variables
(
N {yt , yt−1 , . . . , yt−N , ut−1 , . . . , ut−N } if N > 0
It =
yt if N = 0.

We will construct the Q-value iteration using these information variables. In what follows, we will drop the N depen-
dence on ItN and sometimes we will use N = 1 for simplicity of the notation. For these new approximate states, we
follow the usual Q learning algorithm such that for any I ∈ YN +1 × UN and u ∈ U
 
Qt+1 (I, u) = (1 − αt (I, u))Qt (I, u) + αt (I, u) Ct (I, u) + β min Qt (I1t , v) , (9.21)
v

where I1t = {Yt+1 , yt , . . . , yt−N +1 , ut , . . . , ut−N +1 }, we put the t dependence to emphasize that the distribution of
Yt+1 and hence I1t are different for every t.
To choose the control actions, we use polices that choose the control actions randomly and independent of everything
else such that at time t

ut = ui , w.p σi

for any ui ∈ U with σi > 0 for all i.


We note that for the convergence of the learning algorithm, it is sufficient for the hidden state process to converge to its
invariant distribution under the exploration policy. Hence, any policy that leads the hidden state process to its invariant
measure and visits every action with positive probability can be used for the exploration. For example, the control
action can also be chosen to be a function of the most recent measurement and randomized (as long as all actions have
positive probability of being selected for every measurement realization); this would again lead to a uniquely ergodic
hidden state process under our assumptions. The algorithm is summarized in Algorithm 9.2.1.

Algorithm 9.2.1 Set Parameters: Input: Q0 (initial Q-function), γ ∗ (exploration policy), N (memory window length
for I N ), L (number of data points), {M (I, u)}(I,u) ≡ 0 (number of visits to finite memory state action pairs (I, u)).
Initialize Start with Q0
Iterate If (It , Ut ) is the current memory state-action pair =⇒ generate the cost c(Xt , Ut ) and the next state Yt+1 ∼
T ( · |Xt , Ut ), and update It+1 given It , Yt , Ut .
Iterate set
M (It , Ut ) = M (It , Ut ) + 1.

Iterate For t = 0, . . . , L − 1,
Update Q-function Qt for the inputs (It , Ut ) as follows:

Qt+1 (It , Ut ) = (1 − αt (It , Ut )) Qt (It , Ut )


 
+ αt (It , Ut ) c(Xt , Ut ) + β min Qt (It+1 , v) ,
v∈U

where
1
αt (It , Ut ) = .
1 + M (It , Ut )

43 Generate Ut+1 ∼ γ ∗ .
9.2 Reinforcement Learning Methods for POMDPs 197

End
Return QL

Algorithm 9.2.1 differs from the usual Q-value iterations:


(i) The distribution of I1t , which is the consecutive N-window information variable when we hit the (I, u), is generally
different for every t and the pair (I, u) is not a controlled Markov process.
In other words, the controlled transitions are time dependent, that is, if we assume N = 1 then for some I =
(yt , yt−1 , ut−1 ) and u = ut :

P r(I1t = (yt+1 , yt′ , u′t )|z = (yt , yt−1 , ut−1 ), ut ) = 1{yt =yt′ ,ut =u′t } P r(yt+1 |yt , yt−1 , ut , ut−1 )

is not stationary and might change at every time step t, since P r(yt+1 |yt , yt−1 , ut , ut−1 ) depends on the marginal
distribution of xt−1 (xt−N in the general case).
(ii) Here, we only observe the cost realizations of the underlying state process {xt }t and the control actions. For
example, if we assume that N = 1 then the cost we observe is c(xt , ut ). However, c(xt , ut ) depends on (I, u) pair
randomly and in a time dependent way so that for some I = (yt , yt−1 , ut−1 ) and u = ut :

Ct (I, u) = c(xt , ut ) ∈ B, w.p. P r(Xt ∈ {x : c(x, ut ) ∈ B}|yt , yt−1 , ut−1 )

where P r(dxt |yt , yt−1 , ut−1 ) can be seen as some pseudo-belief on the underlying state variable given I =
(yt , yt−1 , ut−1 ), the most recent N = 1 information variables. In other words, P r(dxt |yt , yt−1 , ut−1 ) is the
Bayesian update of πt−1 , the marginal distribution of the true state xt−1 at the time step t − 1, using I =
(yt , yt−1 , ut−1 ) and thus, it is time dependent. ⋄
We will observe that, if one assumes that the hidden state process, {xt }t is positive Harris recurrent, or at least, admits
a unique invariant probability measure π ∗ under a stationary exploration policy γ, then the average of approximate state
transitions gets closer to

P ∗ (It+1 |It , ut ) := η̂ N ((π ∗ , It+1 )|(π ∗ , It ), ut ) (9.22)

with η̂ N is defined as in (8.25) and (8.26). In particular, if we assume N = 1, then we write



P ∗ (It+1 = (yt+1

, yt′ , u′t )|It = (yt , yt−1 , ut−1 ), ut ) = ⊮{yt′ =yt ,u′t =ut } P π (yt+1 |yt , yt−1 , ut , ut−1 ) (9.23)

where P π (yt+1 |yt , yt−1 , ut , ut−1 ) denotes the distribution of yt+1 when the marginal distribution on xt−1 is given by
the invariant measure π ∗ .
We also have that the sample path averages of the random cost realizations get close to,
Z

∗ ∗
C (I, u) = ĉ(π , I, u) = c(x, u)P π (dx|I)
X

where, P ∗ (x|I) is the Bayesian update of π ∗ , using I. If we assume N = 1, we can write for some I = (y1 , y0 , u0 )
and u = u1
Z

C ∗ (y1 , y0 , u0 , u1 ) = ĉ(π ∗ , (y1 , y0 , u0 ), u1 ) = c(x1 , u1 )P π (dx1 |y1 , y0 , u0 ). (9.24)
X

Now consider the following fixed point equation


X
Q∗ (I, u) = C ∗ (I, u) + β P ∗ (I ′ |I, u) min Q∗ (I ′ , v) (9.25)
v
I′

where P ∗ is defined in (9.22) and C ∗ is defined in (9.24).


198 9 Reinforcement Learning

The existence of a such fixed point follows from usual contraction arguments. The same fixed equation can also be
written as, for N = 1, and for I = (y1 , y0 , u0 ) and u = u1
X ∗
Q∗ ((y1 , y0 , u0 ), u1 ) = C ∗ ((y1 , y0 , u0 ), u1 ) + β P π (y2 |y1 , y0 , u1 , u0 ) min Q∗ ((y2 , y1 , u1 ), v) . (9.26)
v∈U
y2 ∈Y

We note that the stationary distribution π ∗ does not have to be calculated by the decision maker. The Q value iterations
given in (9.21) only use the finite-memory variables I, and π ∗ is not used in the iterations. We will show that the
algorithm naturally converges to (9.25), if the state process is positive Harris recurrent, or at least, admits a unique
invariant probability measure π ∗ under a stationary exploration policy γ, where π ∗ will be the stationary distribution of
the hidden state process xt under the exploration policy. The performance loss will depend on the stationary distribution
π ∗ that is learned via the exploration policy, however, we will establish further upper bounds that are uniform over such
π ∗ which decrease exponentially with the window size N .
That is, one runs Q-learning algorithm by pretending that the finite window is the state. We first need to specify some
conditions that would be needed.

1
Assumption 9.2.1 (i) αk (I, u) = k if Ik = I, uk = u.
(ii) Under the stationary {memoryless or finite memory exploration} policy, say γ, the true state process, {Xt }t , admits
a unique invariant probability measure πγ∗ .
(iii)During the exploration phase, every (I, u) pair is visited infinitely often.

Condition (iii) can be relaxed: However, one needs to ensure that the set of all (y, u) pairs which are visited infinitely
often during exploration is so that an optimal policy is learned (visited infinitely often), and when this optimal policy
(learned via the convergence of Q-learning) is implemented, the closed-loop process always remains in this set; see
[190] and [91, Lemma 6 and Corollary 2].

Theorem 9.2.1 [188] Under the previous assumption, the algorithm given by
 
Qk+1 (I, u) = (1 − αk (I, u))Qk (I, u) + αk (I, u) Ck (I, u) + β min Qk (I1 , v) ,
v

converges almost surely to Q∗ which satisfies


X
Q∗ (I, u) = C ∗ (I, u) + β P ∗ (I ′ |I, u) min Q∗ (I ′ , v)
v
I′

which are the Q values for the approximate belief MDP.

Corollary 9.2.1 [188] Under the filter stability conditions, finite window Q learning is nearly optimal:
(a) Under the exponential filter stability condition, for any policy γ N that satisfies Q∗ (I, γ N (I)) = minu Q∗ (I, u)

− − 4∥c∥∞ N
, T , γ N ) − Jβ∗ (πN , T )|I0N ≤
 
E Jβ (πN α .
(1 − β)2

(b) If we only have asymptotic filter stability (uniform in policies) in total variation, as N → ∞,
− −
, T , γ N ) − Jβ∗ (πN , T )|I0N → 0.
 
E Jβ (πN

Remark 9.2. Under the conditions of Corollary 8.6, the bounds via the Hilbert metric provide complementary sufficient
conditions for geometric decay of the approximation error to zero in the memory length.

Remark 9.3. We caution the reader that our result assumes that the cost starts running after time N : that is the effective
cost is:
9.3 Q-Learning For Continuous State and Action Spaces: Quantized Q-Learning, its Convergence and Near-Optimality 199

" #
X
k−N
E β c(xk , uk ) . (9.27)
k=N

Of course, this criterion is also applicable if the system starts running prior to time −N and the costs become in effect
after time 0.
If this criterion is not applicable, and the first N stages are also crucial, (i) if β is large enough, we can conclude that
the first N stages are not as critical for the analysis as their contributions will be minor in comparison with the future
stages for the criterion, which can also be seen by noting that for large enough β, the contributions of the first N time
stages become negligible: "∞ #
X
k
(1 − β)E β c(xk , uk ) .
k=0

(ii) On the other hand, if β is not large and if the cost starts running at time 0, then, we can first run the Q-learning
algorithm above to find the best N -window policies which optimizes (9.27). The remaining question would be to
optimize:
"N −1 #
X
E c(xk , uk ) + V (Ik ) (9.28)
k=0

as a finite-horizon optimal control problem with a terminal cost and the terminal cost V can be estimated by (9.27) via
Theorem 8.3.1 and Theorem 9.2.1. The question then becomes how to select the first N actions, leading to a problem
with a finite search complexity for a finite horizon problem, without knowing the system dynamics. For this, one can
run a MCMC algorithm in parallel simulations to find the optimal policy for the first N time stages. Since the resulting
policy minimizing (9.28) will be at least as good as the first N -window policy under the optimal (belief-MDP) policy
(which is not designed to optimize (9.28) but the original cost, the bounds presented in Theorem 8.3.2 will be applicable
even when the cost criterion includes the first N time stages.

See Theorem 6.4.4 for sufficient conditions for asymptotic filter stability under total variation.

9.3 Q-Learning For Continuous State and Action Spaces: Quantized Q-Learning, its
Convergence and Near-Optimality

With the approach above, by quantizing the state space and viewing the quantization output as a measurement (and
quantizer as a measurement kernel), and thus as a POMDP; we can arrive at an approximate finite MDP. First note that
under Assumption 8.2.1, Theorem 8.2.1 leads to near-optimality of finite actions.
Let ρ ≡ Qm , be a nearest-neighbour quantizer with finite range Xm :

Qm (z) := arg min dX (z, zk ).


zk ∈Xm

Then, one runs the following: Using the nearest neighbour map ρ, write for any (x, u) ∈ X × U

Qk+1 (ρ(x), u) = (1 − αk (ρ(x), u))Qk (ρ(x), u) (9.29)


 
+ αk (ρ(x), u) Ck (ρ(x), u) + β min Qk (ρ(X1 ), v)
v

that is for any true value of the state, we use its representative state from the finite set Xm : Thus, one run Q-learning
as if the quantized state is the actual state.
200 9 Reinforcement Learning

For exploration, we again use polices that choose the control actions randomly and independent of everything else with
positive probability for every action: the invariant measure of the state process xt (under the exploration policy) should
put positive measure on these bins and satisfy an ergodicity condition to be presented in the following:

Assumption 9.3.1 (i) With y = ρ(x), we let αt (y, u) = 0 unless (Yt , Ut ) = (y, u). Otherwise, let
1
αt (y, u) = Pt .
1+ 1
k=0 {Yk =I,Uk =u}

(ii) Under the exploration policy γ ∗ , the state process is uniquely ergodic (and thus has a unique invariant probability
measure πγ ∗ .
(iii)During the exploration phase, every observation-action pair (y, u) is visited infinitely often.

As noted in Remark 4.3, Meyn and Tweedie [232, Theorem 13.0.1] show that for an aperiodic Harris recurrent Markov
chain, for each initial state x ∈ X,
lim sup |P n (x, B) − π(B)| = 0,
n→∞ B∈B(X)

that is P n (x, ·) converges to π in total variation. This assumption is sufficient, but not necessary for the second item to
hold.

Algorithm 9.3.1 Set Parameters: Input: Q0 (initial Q-function) q : X → Y (quantizer), γ ∗ (exploration policy), L
(number of data points), {N (y, u) = 0}(y,u)∈×U (number of visits to state-action pairs).
Initialize Start with Q0
Iterate If (Xt , Ut ) is the current state-action pair =⇒ generate the cost c(Xt , Ut ) and the next state Xt+1 ∼
T ( · |Xt , Ut ),
Iterate set
N (q(Xt ), Ut ) = N (q(Xt ), Ut ) + 1.

Iterate For t = 0, . . . , L − 1,
Update Q-function Qt for the inputs (q(Xt ), Ut ) as follows:

Qt+1 (q(Xt ), Ut ) = (1 − αt (q(Xt ), Ut )) Qt (q(Xt ), Ut )


 
+ αt (q(Xt ), Ut ) c(Xt , Ut ) + β min Qt (q(Xt+1 ), v) ,
v∈U

where
1
αt (q(Xt ), Ut ) = .
1 + N (q(Xt ), Ut )

43 Generate Ut+1 ∼ γ ∗ .
End
Return QL

Theorem 9.3.1 [184] Under Assumption 9.2.1, for every pair (yi , u) ∈ Y × U, the algorithm given above converges
to
X
Q∗ (yi , u) = C ∗ (yi , u) + β P ∗ (yj |yi , u) min Q∗ (yj , v).
v∈U
yj ∈Y
9.3 Q-Learning For Continuous State and Action Spaces: Quantized Q-Learning, its Convergence and Near-Optimality 201

Here, P ∗ and C ∗ are defined by


Z

C (yi , u) = c(x, u) π̂y∗i (dx)
Bi
Z

P (yj |yi , u) = T (Bj |x, u) π̂y∗i (dx), (9.30)
Bi

where
πγ ∗ (A)
π̂y∗i (A) := , ∀A ⊂ Bi , ∀i ∈ {1, . . . , M }, (9.31)
πγ ∗ (Bi )

and πγ ∗ is the invariant measure of the state process under the exploration policy γ ∗ .

Remark 9.4. Observe the similarity with (8.18). Here, however, the weighting measures π̂y∗i are not arbitrarily selected,
and are induced by the exploration policies.

9.3.1 Error Analysis for Convergence of Quantized Q-Learning for Continuous Space MDPs

The model described in (8.18) is the same model given by the equations (9.30). Hence, the results and error bounds
from Section 8.2.2 can be used for the error analysis of the Q-learning algorithm given by (9.3.1). For the remainder
of this section, we present a series of results for the performance of the policies learned through the approximate Q-
learning algorithm in (9.3.1) building on the results from Section 8.2.2. In these corollaries, it is always assumed that
πγ ∗ (Bi ) > 0 for all i ∈ {1, . . . , M } where Bi ’s are the quantization bins and πγ ∗ is the invariant measure on the state
process under the exploration policy γ ∗ .

Error Analysis for Non-Compact MDPs

The first result is in asymptotic nature and requires very mild conditions for the convergence (i.e., continuity of the
stage-wise cost and weak continuity of the transition kernel). It follows from Theorem 9.3.1 and Theorem 8.2.9.

Corollary 9.3.1 [184] Under Assumption 9.2.1 and Assumption 8.2.1, the Q learning algorithm in (9.3.1) converges
to Q∗ in Theorem 9.3.1 with probability 1 and for any policy γ̂ that satisfies Q∗ (x, γ̂(x)) = minu∈U Q∗ (x, u) (i.e.,
greedy policy of Q∗ ), for any compact K ⊂ X, we have

sup Jβ (x0 , γ̂) − Jβ∗ (x0 ) → 0


x0 ∈K

as L− → 0, where L− is defined in (8.22).

We recall now that the error bounds to be presented in Corollary 9.3.2 and Corollary 9.3.3 below will involve the
function L and the uniform bound L̄ which are defined as follows: for some x ∈ X where x belongs to a quantization
bin Bi whose representative state is yi (i.e. q(x) = yi ) and averaging measure π̂y∗i , we have
Z
L(x) := ∥x − x′ ∥ π̂y∗i (dx′ )
Bi
L̄ := max sup ∥x − x′ ∥.
i=1,...,M x,x′ ∈Bi

The following result follows from Theorem 9.3.1 and Theorem 8.2.6.
202 9 Reinforcement Learning

Corollary 9.3.2 [184] Under Assumption 9.2.1 and Assumption 8.2.4, the Q-learning algorithm in (9.3.1) converges
to Q∗ in Theorem 9.3.1 with probability 1 and for any policy γ̂ that satisfies Q∗ (x, γ̂(x)) = minu∈U Q∗ (x, u) (i.e.,
greedy policy of Q∗ ), for any initial state x0 , we have
  ∞
βαT ∥c∥∞ X t
Jβ (x0 , γ̂) − Jβ∗ (x0 ) ≤ 2 αc + β sup Exγ0 [L(Xt )] .
1−β t=0 γ∈Γ

Application to Models with Compact State Spaces

For the case with compact spaces, we obtain sharper bounds in the following.
The following result follows from Theorem 9.3.1 and Theorem 8.2.8.

Corollary 9.3.3 [184] Under Assumption 9.2.1 and Assumption 8.2.5, the Q learning algorithm in (9.3.1) converges
to Q∗ in Theorem 9.3.1 with probability 1 and for any policy γ̂ that satisfies Q∗ (x, γ̂(x)) = minu∈U Q∗ (x, u) (i.e.,
greedy policy of Q∗ ), we have
2αc
sup Jβ (x0 , γ̂) − Jβ∗ (x0 ) ≤ 2
L̄.
x0 ∈X (1 − β) (1 − βαT )

where L̄ is defined in (8.21).

Building on the results presented, we now show that for compact state spaces, the terms L(x) and the uniform bound L̄
can be explicitly bounded via cardinality of finite approximating set Y and dimension d of the state space. To this end,
we assume that the state space X ⊂ Rd is compact, and thus totally bounded. Then, for a given M , we can quantize X
by choosing a finite subset Y = {y1 , . . . , yM } such that

max min ∥x − yi ∥ ≤ α(1/M )1/d


x∈X yi ∈Y

for some α > 0, which is possible since X is totally bounded ( [110, Theorem 2.3.1]). Using this construction, one can
then write the following immediate bounds:

L(x) ≤ 2α(1/M )1/d , for all x ∈ X,


L̄ ≤ 2α(1/M )1/d .

We can then state the following results, which follow from Corollary 9.3.2 and Corollary 9.3.3.

Corollary 9.3.4 [184] If the state space X ⊂ Rd is compact, under Assumption 9.2.1 and Assumption 8.2.4, the Q-
learning algorithm in (9.3.1) converges to Q∗ in Theorem 9.3.1 with probability 1 and for any policy γ̂ that satisfies
Q∗ (x, γ̂(x)) = minu∈U Q∗ (x, u) (i.e., greedy policy of Q∗ ), for any initial state x0 , we have

4α(1/M )1/d
 
βαT ∥c∥∞
Jβ (x0 , γ̂) − Jβ∗ (x0 ) ≤ αc +
1−β 1−β

Corollary 9.3.5 [184] If the state space X ⊂ Rd is compact, under Assumption 9.2.1 and Assumption 8.2.5, the Q
learning algorithm in (9.3.1) converges to Q∗ in Theorem 9.3.1 with probability 1 and for any policy γ̂ that satisfies
Q∗ (x, γ̂(x)) = minu∈U Q∗ (x, u) (i.e., greedy policy of Q∗ ), we have
4αc
sup Jβ (x0 , γ̂) − Jβ∗ (x0 ) ≤ α(1/M )1/d
x0 ∈X (1 − β)2 (1 − βαT )
9.4 A General Q-Learning Convergence Theorem 203

9.4 A General Q-Learning Convergence Theorem

In this section, we present a generalization of the convergence results presented earlier in the chapter. In many problems
including most of those in applied, health, and social sciences, and financial mathematics, one may not even know
whether the problem studied can be formulated as a fully observed Markov Decision Process (MDP), or a partially
observable Markov Decision Process (POMDP) or a multi-agent system where other agents are present or not. There
are many practical settings where one works with data and does not know the possibly very complex structure under
which the data is generated and tries to respond to the environment. For such settings, a common practical approach is
to view the system as an MDP, with a perceived state and action (which may or may not define a genuine controlled
Markov chain and therefore, the MDP assumption may not hold in actuality), and arrive at corresponding solutions via
some learning algorithm.
Toward this end, a general convergence theorem was given in [190, Theorem 2.1], with further implications and refine-
ments reported in [91, Section IV.B]. This is presented in the following:
Let {Ct }t be R-valued, {St }t be S-valued and {Ut }t be U-valued three stochastic processes. Consider the following
iteration defined for each (s, u) ∈ S × U pair

Qt+1 (s, u) = (1 − αt (s, u)) Qt (s, u) + αt (s, u) (Ct + βVt (St+1 )) (9.32)

where Vt (s) = minu∈U Qt (s, u), and αt (s, u) is a sequence of constants also called the learning rates. We note that un-
like the finite setup considered earlier where wt ∈ (9.8)wasaconditionallyzero−meanmartingalenoise, suchapropertydoesnotapp
M arkoviannatureof thedynamics.
We assume that the process Ut is selected so that the following condition holds.

Assumption 9.4.1 S, U are finite sets, and the joint process (St+1 , St , Ut , Ct )t≥0 is asymptotically ergodic in the sense
that for the given initialization random variable S0 , for any measurable bounded function f , we have that with probability
one,
N −1 Z
1 X
f (St+1 , St , Ut , Ct ) → f (s1 , s, u, c)π(ds1 , ds, du, dc)
N t=0

for some measure π such that π(S × s × u × R) > 0 for any (s, u) ∈ S × U.

Remark 9.5. The assumption that π(S × s × u × R) > 0 for any (s, u) ∈ S × U is in the same spirit as the standard
condition for reinforcement algorithms that every state-action pair is visited infinitely often during training. We note that
it is possible to relax this condition, if one is only interested in the convergence of the algorithm. In particular, we might
consider a measure π such that π(S × s × u × R) > 0 for all (s, u) ∈ B ⊂ S × U, for some subset B, where the set B
represents so called trained state-action pairs. For the learned policies to be optimal, however, one needs to make sure that
the controlled process stays within the trained part of the system during the execution of an optimal policy.

The above implies Assumption 9.4.2(ii)-(iii) below:

Assumption 9.4.2 i. αt (s, u) = 0 unless (St , Ut ) = (s, u). Furthermore,


1
αt (s, u) = Pt
1+ k=0 1{Sk =s,Uk =u}
P
and with probability 1, t αt (s, u) = ∞
ii. For Ct , we have, as t → ∞,
Pt
Ck 1{Sk =s,Uk =u}
k=0
Pt → C ∗ (s, u),
k=0 1{Sk =s,Uk =u}
204 9 Reinforcement Learning

almost surely for some function C ∗ .


iii. For the St process, we have, for any function f , as t → ∞,
Pt
k=0 f (Sk+1 )1{Sk =s,Uk =u}
Z
Pt → f (s1 )P ∗ (ds1 |s, u)
k=0 1{Sk =s,Uk =u}

almost surely for some P ∗ .

Note that a stationarity assumption is not required. Under Assumption 9.4.1, we have that with f (St+1 , St , Ut , Ct ) =
Ct 1{St =s,Ut =u} , as N → ∞,

N −1 Z
1 X
Ct 1{St =s,Ut =u} → Cπ(S = s, U = u, dC).
N t=0 C∈R

We also have that with f (St+1 , S1 , Ut , Ct ) = 1{St =s,Ut =u} , as N → ∞,

N −1
1 X
1{St =s,Ut =u} → π(S = s, U = u)
N t=0

almost surely. Hence, we can write


1
Pt
Ck 1{Sk =s,Uk =u}
Z
t+1 k=0
1
Pt → Cπ(dC|S = s, U = u) =: C ∗ (s, u)
t+1 k=0 1{Sk =s,Uk =u}

which implies Assumption 9.4.2 (ii). Similarly, one can also establish Assumption 9.4.2 (iii) under Assumption 9.4.1.
As before, let S, U be finite sets. Consider the following equation
X
Q∗ (s, u) = C ∗ (s, u) + β V ∗ (s1 )P ∗ (s1 |s, u) (9.33)
s1 ∈S

for some functions Q∗ , C ∗ , to be defined explicitly, and for some regular conditional probability distribution P ∗ (·|s, u),
where V ∗ (u) := minu Q∗ (s, u).

Theorem 9.4.1 [190, Theorem 2.1] Under Assumption 9.4.2, Qt (s, u) → Q∗ (s, u) almost surely for each (s, u) ∈ S × U
pair where Q∗ satisfies (9.33), for any initialization of Q0 .

We note that this result generalizes those presented earlier in this chapter and in particular Theorem 9.2.1. See also [77,
106, 188].

9.5 Bibliographic Notes

Q-learning was introduced and studied in [323], [306], [26]. An ODE approach to Q-learning presents a rather direct proof
of convergence [64] [319]. Two comprehensive resources on reinforcement learning are [301] and [230].
Quantized Q-learning and its convergence and near-optimality is studied in [184].
Q-learning for partially observed MDPs have been studied in [137, 228, 289, 302]. Our analysis here builds on [188] which
also establishes convergence to near optimality.
9.6 Exercises 205

9.6 Exercises

Exercise 9.6.1 Consider a controlled Markov chain with state space X = {0, 1}, action space U = {0, 1}, and transition
kernel for t ∈ Z+ :
P (xt+1 = 1|xt = 0, ut = 1) = P (xt+1 = 1|xt = 1, ut = 1) = α
P (xt+1 = 1|xt = 0, ut = 0) = P (xt+1 = 1|xt = 1, ut = 0) = 1 − α.
where α ∈ (0, 1). Let a cost function c(x, u), with c : X × U → R+ be given by

c(0, 1) = c(0, 0) = 1 c(1, 0) = c(1, 1) = 2.

Suppose that the goal is to minimize the quantity



X
E0γ [ β t c(xt , ut )],
t=0

for a fixed β ∈ (0, 1), over all admissible policies γ ∈ ΓA .


Find an optimal policy and the optimal expected cost explicitly, as a function of α, β (note that the initial condition is
x0 = 0).

Exercise 9.6.2 Consider the following problem: Let X = {1, 2}, U = {1, 2}, where X denotes whether a fading channel is
in a good state (x = 2) or a bad state (x = 1). There exists an encoder who can either try to use the channel (u = 2) or
not use the channel (u = 1). The goal of the encoder is send information across the channel.
Suppose that the encoder’s cost (to be minimized) is given by:

c(x, u) = −1{x=2,u=2} + α(u − 1),

for α = 1/2 (if you view this as a maximization problem, you can see that the goal is to maximize information transmission
efficiency subject to a cost involving an attempt to use the channel; the model can be made more complicated but the idea
is that when the channel state is good, u = 2 can represent a channel input which contains data to be transmitted and
u = 1 denotes that the channel is not used).
Suppose that the transition kernel is given by:

P (xt+1 = 2|xt = 2, ut = 2) = 0.8, P (xt+1 = 1|xt = 2, ut = 2) = 0.2

P (xt+1 = 2|xt = 2, ut = 1) = 0.2, P (xt+1 = 1|xt = 2, ut = 1) = 0.8


P (xt+1 = 2|xt = 1, ut = 2) = 0.5, P (xt+1 = 1|xt = 1, ut = 2) = 0.5
P (xt+1 = 2|xt = 1, ut = 1) = 0.9, P (xt+1 = 1|xt = 1, ut = 1) = 0.1

We will consider either a discounted cost criterion for some β ∈ (0, 1) (you can fix an arbitrary value)

X
inf Exγ [ β t c(xt , ut )] (9.34)
γ
t=0

or the average cost criterion


T −1
1 γ X
inf lim sup Ex [ c(xt , ut )]. (9.35)
γ T →∞ T t=0

a) Using Matlab or some other program, obtain a solution to the problem given above in (9.34) through the following:
206 9 Reinforcement Learning

(i) Policy Iteration


(ii) Value Iteration.
(iii)Q-Learning. Note that a common way to pick α coefficients in the Q-learning algorithm is to take for every x, u
pair:
1
αt (x, u) = Pt
1 + k=0 1{xk =x,uk =u}

b) Consider the criterion given in (9.35). Apply the convex analytic method, by solving the corresponding linear program,
to find the optimal policy. In Matlab, the command linprog can be used to solve linear programming problems. See (7.42).

Exercise 9.6.3 Revisit Exercise 9.6.3, part a). Apply Q-Learning, noting that a common approach to pick α coefficients in
the Q-learning algorithm is to take for every x, u pair:
1
αt (x, u) = Pt
1+ k=0 1{xk =x,uk =u}

Exercise 9.6.4 One can apply Q-learning even when the model is not known, but for the results to be optimal, it is essential
that the system we are dealing with is an actual MDP.
a) However, imagine that we have a POMDP but we run the Q-learning algorithm as if the system is an MDP. Sections 9.2
and 9.3 have shown that Q-learning, when a finite window memory of the most recent measurements and actions is viewed
as a state, converges even in this case, and even to near optimality under mild conditions related to either filter stability or
appropriate approximation bounds. Study Exercise 7.7.6 and apply finite memory Q-learning for this example.
b) [Quantized Q-learning]As a further instance we have considered the finite-state quantization of a continuous space
MDP and view this as a POMDP.
Consider the setup in Exercise 9.6.3. Revise the problem above with the following transition kernel so that X = [0, 1] (thus
the channel’s quality is not binary) and for each Borel A ∈ [0, 1]
Z Z
P (xt+1 ∈ A|xt = z0 , ut = 1) = 2 (1 − x)dx, P (xt+1 ∈ A|xt = z0 , ut = 0) = 2 xdx,
A A

for all z0 ∈ [0, 1], and suppose that the encoder’s per-stage cost (to be minimized) is given by:

c(x, u) = −xu + ηu.

for some η ∈ R. Apply quantized Q-learning by quantizing the channel state with uniform quantization of increasing
granularity.

Exercise 9.6.5 Recall Exercise 7.7.6: Consider a POMDP given by the following description. Let there be two possible
states that a machine can take: X = {0, 1}, where 0 is the bad (‘system is down’) state and 1 is the good state. Let
U = {0, 1}, where 0 is the ‘do nothing’ control and 1 is the ‘repair’ control. Suppose that the transition probabilities are
given by:

P (Xt+1 = 1|Xt = 1, Ut = 0) = 1 − η1 , P (Xt+1 = 0|Xt = 1, Ut = 0) = η1 > 0


P (Xt+1 = 1|Xt = 1, Ut = 1) = 1 − η2 , P (Xt+1 = 0|Xt = 1, Ut = 1) = η2 > 0
P (Xt+1 = 1|Xt = 0, Ut = 0) = 0, P (Xt+1 = 0|Xt = 0, Ut = 0) = 1
P (Xt+1 = 1|Xt = 0, Ut = 1) = α > 0, P (Xt+1 = 0|Xt = 0, Ut = 1) = 1 − α (9.36)

Thus, η1 is the failure probability when the state is good (and no repair) and η2 is the failure probability when the state is
good (and when there is repair) with η1 > η2 , and α is the success probability in the event of a repair.
The controller has access only to {0, 1}-valued measurement variables Y0 , · · · , Yt and U0 , · · · , Ut−1 , at time t, where the
measurements are generated by a binary symmetric channel:
9.6 Exercises 207

(Y = x|X = x) = 1 − ϵ, P (Y = 1 − x|X = x) = ϵ,

for all x ∈ {0,. The per-stage cost function c(x, u) is given by c(0, 0) = C, c(1, 0) = 0, c(0, 1) = c(1, 1) = R with
0 < R < C.
Apply Q-learning using the finite window ItN = {[yt−N , t]}, [ut−N , t − 1]}} as an effective state, as described in Section
9.2 of the lecture notes. When can you guarantee convergence to near-optimality as the window size increases? Compute
the numerical performance for several N values. Reflect on the trade-off between performance, computational demands,
and memory length for such a control design.
10

Decentralized and Multi-Agent Stochastic Control

10.1 Introduction

In classical stochastic control problems considered so far in this document, we were given a system of the form

xt+1 = f (xt , ut , wt ), t ∈ Z+ ,

where actions are to be generated using some control policy γ = {γt } with

ut = γt (It ), t ∈ Z+ ,

where It is the information available at t. Here, wt is i.i.d. noise. If It = {x0 , · · · , xt ; u0 , · · · , ut−1 }, we have a fully
observed system. If the controller only has measurements

yt = g(xt , vt ),

It = {y0 , · · · , yt ; u0 , · · · , ut−1 }, we have a partially observed system. These have been studied extensively in the previous
chapters.
As we observed earlier in these notes, given an optimality criterion (e.g. expected finite horizon cost, discounted cost,
average cost, terminal cost), for such classical stochastic control setups, there are few powerful techniques to establish the
existence/computation of optimal policies:
(i) The dynamic programming approach and backward induction: Weak-continuity / strong continuity properties and
measurable selection conditions leads to existence / explicit computations.
(ii) The strategic measures approach (see Section 5.4).
(iii) For infinite horizon problems, linear programming/convex analytic techniques.
All of these crucially build on the fact that It ⊂ It+1 , that is, information is expanding. In the absence of this condition,
which facilitates the applicability of the iterated expectations theorem (Theorem 4.1.3), much of the standard analysis on
existence/structure/recursions is no longer applicable: The reader is referred to the derivation at the beginning of Chapter
5.
However, a very important class of optimal stochastic control problems involve setups where a number of decentralized
decision makers are present. In this context, we will consider a collection of decision makers (DMs) where each has access
to some local information variable: Such a collection of decision makers who wish to minimize a common cost function
and who has an agreement on the system (that is, the probability space on which the system is defined, and the policy and
action spaces) is said to be a stochastic team. Such problems are called decentralized stochastic control problems.
To gain some insight, let us consider the following model.

xt+1 = f (xt , u1t , . . . , uL


t , wt ),
210 10 Decentralized and Multi-Agent Stochastic Control

with each decision maker DM m arriving at their action um at time t using only local information:

um m m
t = γt (It ), t ∈ Z+ ,

where Itm denotes some information variable. Decentralized stochastic control theory requires more general approaches
when compared with the classical setup that we have considered up until this chapter, primarily due to the informational
subtleties, to be presented further in the following.
To study such problems in a systematic fashion, we will present a classification for decentralized stochastic control models
based on the informational and dynamical relations between the decision makers in the following. Toward this goal, in the
following we introduce Witsenhausen’s intrinsic model.

10.2 Solution Concepts, Information Structures and Witsenhausen’s Intrinsic Model

10.2.1 Witsenhausen’s intrinsic model

Witsenhausen’s contributions (e.g., [327, 328, 331]) to decentralized stochastic control and characterization of information
structures have been crucial in our understanding of stochastic team theory. In this section, we introduce the character-
izations as laid out by Witsenhausen, termed as the Intrinsic Model [328]; see [349] for a comprehensive overview and
further characterizations and classifications of information structures. In this model (described in discrete time), any action
applied at any given time is regarded as applied by an individual decision maker/agent, who acts only once. One advantage
of this model, in addition to its generality, is that the characterizations regarding information structures can be compactly
described.
Suppose that in the decentralized system considered below, there is a pre-defined order in which the decision makers act.
Such systems are called sequential systems (for non-sequential teams, we refer the reader to Andersland and Teneketzis
[8], [9] and Teneketzis [304], in addition to Witsenhausen [326] and [349, p. 113]). Suppose that in the following, the
action and measurement spaces are standard Borel spaces, that is, Borel subsets of Polish (complete, separable and metric)
spaces. In the context of a sequential system, the Intrinsic Model of Witsenhausen [329] is the following characterization
of information structures, where we consider a decentralized stochastic control model with N decision makers (DMs) (also
called agents). In the model, there exist the following:
– A collection of measurable spaces {(Ω, F), (Ui , U i ), (Yi , Y i ), i ∈ N }, specifying the system’s distinguishable
events, and the control and measurement spaces. Here N = |N | is the number of control actions taken, and each
of these actions is taken by an individual (different) DM (hence, even a DM with perfect recall can be regarded
as a separate decision maker every time it acts). The pair (Ω, F) is a measurable space (on which an underlying
probability may be defined). The pair (Ui , U i ) denotes the measurable space from which the action, ui , of decision
maker i is selected. The pair (Yi , Y i ) denotes the measurable observation/measurement space for DM i.
– A measurement constraint which establishes the connection between the observation variables and the system’s
distinguishable events. The Yi -valued observation variables are given by y i = η i (ω, u[1,i−1] ), u[1,i−1] = {uk , k ≤
i − 1}, η i measurable functions and uk denotes the action of DM k. Hence, the information variable y i induces a σ-
Qi−1
field, σ(I i ) over Ω × k=1 Uk . The collection {J i ; i = 1, . . . , N } or {η i ; i = 1, . . . , N } is called the information
structure of the system.
– A design constraint which restricts the set of admissible N -tuple control laws γ = {γ 1 , γ 2 , . . . , γ N }, also called
[1,i−1]
designs or policies, to the set of all measurable control functions, so that ui = γ i (y i ), with y i = η i (ω,
Qu ),
and γ i , η i measurable functions. Let Γ i denote the set of all admissible policies for DM i and let Γ = k Γ k .
We note that, the intrinsic model of Witsenhausen gives a set-theoretic characterization of information fields, however, for
standard Borel spaces, the model above is equivalent to that of Witsenhausen’s; see Exercise 1.6.6.
One can also introduce a fourth component.
– A probability measure P defined on (Ω, F) which describes the measures on the events in the model.
10.2 Solution Concepts, Information Structures and Witsenhausen’s Intrinsic Model 211

10.2.2 Solution concepts

Thus, we will assume that we are given a probability measure P on (Ω, F). Additionally, we have a loss (or cost) function
c : Ω0 × (U1 × · · · × UN ) → R+ to be optimized where Ω0 is an appropriate signal space.
Let
γ = {γ 1 , · · · , γ N } ∈ Γ.
We then have,

J(γ) = E[c(ω0 , u)] = E[c(ω0 , γ 1 (y 1 ), · · · , γ N (y N ))], (10.1)

for some non-negative measurable loss (or cost) function c : Ω × k Uk → R+ . Here, we have the notation u = {ut , t ∈
Q
N }. Here, ω0 may be viewed as the cost function relevant exogenous variable and is contained in ω.

Definition 10.2.1 For a given stochastic team problem with a given information structure, {J; Γ i , i ∈ N }, a policy (strat-
∗ ∗
egy) N -tuple γ ∗ := (γ 1 , . . . , γ N ) ∈ Γ is an optimal team decision rule (team-optimal decision rule or simply team-
optimal solution) if
J(γ ∗ ) = inf J(γ) =: J ∗ , (10.2)
γ∈Γ

provided that such a strategy exists. The cost level achieved by this strategy, J ∗ , is the minimum (or optimal) team cost.

Definition 10.2.2 For a given N -person stochastic team with a fixed information structure, {J; Γ i , i ∈ N }, an N -tuple of
∗ ∗
strategies γ ∗ := (γ 1 , . . . , γ N ) constitutes a Nash equilibrium (synonymously, a person-by-person optimal (pbp optimal)
solution) if, for all β ∈ Γ i and all i ∈ N , the following inequalities hold:

J ∗ := J(γ ∗ ) ≤ J(γ −i∗ , β), (10.3)

where we have adopted the notation


∗ ∗ ∗ ∗
(γ −i∗ , β) := (γ 1 , . . . , γ i−1 , β, γ i+1 , . . . , γ N ). (10.4)

 
For notational simplicity, let for any 1 ≤ k ≤ N , γ −k := γ i , i ∈ {1, · · · , N } \ {k} . In the following, we will denote
by bold letters the ensemble of random variables across the DMs; that is, y = {y i , i = 1, · · · , N } and u = {ui , i =
1, · · · , N }.

Example 10.1. Consider the following model of a system with two decision makers [349]. Let Ω = {ω1 , ω2 , ω3 }, F be
the power set of Ω. Let the action space be U1 = {U (up), D(down)}, U2 = {L(left), R(right)}, and U 1 and U 2 be the
power sets of U1 and U 2 respectively.
Suppose the probability measure P is given by P (ωi ) = pi , i = 1, 2, 3 and p1 = p2 = 0.3, p3 = 0.4, and the loss function
c(ω, u1 , u2 ) is given by the following matrices

u2 u2 u2
L R L R L R
u1 U 1 0 U 2 3 U 1 2
D 3 1 D 2 1 D 0 2
ω : ω1 ↔ 0.3 ω2 ↔ 0.3 ω3 ↔ 0.4

Case 1. First, let us consider the case where both agents have access to the true state of nature, and hence Y1 = Y2 =
σ({{ω1 }, {ω2 }, {ω3 }}), the σ−field generated by the singletons.
In this case, the unique team-optimal decision rules are:
212 10 Decentralized and Multi-Agent Stochastic Control
 
 U, ω = ω1  L, ω = ω3
∗ ∗
γ 1 (ω) = γ 2 (ω) = ,
D, else R, else
 

which we may write symbolically as


γ ∗ = (U DD, RRL).
We point to the observation that even though the policy pair (U DD, RRL) is unique as a team-optimal solution (which
is also, by definition, pbp optimal), it is not the unique pbp optimal solution. The policy pair (U U D, RLL) is also pbp
optimal, but it is suboptimal.
Case 2. Let the information fields J 1 = {∅, {ω1 }, {ω2 , ω3 }, Ω} and J 2 = {∅, {ω1 , ω2 }, {ω3 }, Ω}.
For the above model, the unique optimal control strategy is given by
(
1,∗ 1 U, y 1 = {ω1 }
γ (y ) =
D, else
(
R, y 2 = {ω1 , ω2 }
γ 2,∗ (y 2 ) =
L, else

The development of a systematic solution approach in optimal decentralized stochastic control requires a cautious classifi-
cation of such problems, primarily in view of information structures.

10.2.3 Classification of information structures

Static vs. dynamic information structures


Under the intrinsic model presented, an Information structure (IS) is dynamic if the information available to at least one
DM is affected by the action of at least one other DM. An IS is static, if the information available at every decision maker
is only affected by exogenous disturbances:
(i) A sequential team is static, if the information available at every decision maker is only affected by exogenous
disturbances (Nature); that is no other decision maker can affect the information at any given decision maker.
(ii) A sequential team problem is dynamic if the information available to at least one DM is affected by the action of at
least one other DM.
Figure 10.1 is a depiction for a static team problem.

ω0

Q1 Q2 Q3

y1 y2 y3

γ1 γ2 γ3

u2 u3
u1

Fig. 10.1: An example of a static information structure. Here, Qi (y i ∈ ·|ω0 ) := P (η i (ω) ∈ ·|ω0 ), i = 1, 2, 3.

Classical, quasi-classical (partially nested), and nonclassical information structures


10.2 Solution Concepts, Information Structures and Witsenhausen’s Intrinsic Model 213

(i) An IS {y i , 1 ≤ i ≤ N } is classical if y i contains all of the information available to DM k for k < i. E.g.: Classical
stochastic control problems; the material in these lecture notes up to this chapter (involving fully observed or
partially observed models) has thus been with regard to classical information structures.
(ii) An IS is quasi-classical or partially nested, if whenever uk , for some k < i, affects y i , y i contains y k .
(iii)An IS which is not partially nested is nonclassical.

Fully and partially observable models reviewed earlier in the chapter are classical in the sense described above. Non-
classical problems can be very challenging. As we will see in Section 10.4.2, even a linear system with Gaussian noise can
lead to a problem which is very difficult to study and can admit an optimal solution which is not linear; this example is
known as Witsenhausen’s counterexample [325]. This is called a counterexample since it shows that even linear quadratic
Gaussian (LQG) models can admit solutions which are not linear and hence it is a counterexample to a natural conjecture
(that all solutions to such LQG problems should be linear) when the information structure is nonclassical.
Quasi-classical information structures possess useful characteristics which allow for solution methods tailored for such
models, as we will see in the chapter.

10.2.4 A state space model

A sub-class of sequential teams involve setups where there is a controlled state model, where unlike the classical single-
agent model, the state realizations are not available to the agents. In a state space model, one assumes that the decentralized
control system has a state xt that is evolving with time. The evolution of the state is controlled by the actions of the agents
(control stations). We may assume that the system has N control stations where each control station i chooses a control
action uit at time t. The system considered runs in discrete time, either for a finite or an infinite horizon. In the context of
Witsenhausen’s intrinsic model, any decision maker applying an action at a given time stage is interpreted as a different
decision maker.
Let X denote the space of realizations of the state xt , and Ui denote the space of realization of control actions uit . Let T
denote the set of time for which the system runs.
The initial state x1 is a random variable and the state of the system evolves as

xt+1 = ft (xt , u1t , . . . , uN 0


t ; wt ) , t∈T, (10.5)

where {wt0 , t ∈ T } is an independent noise process that is also independent of x1 . We assume that each control station i
observes the following at time t
yti = gti (xt , wti ), (10.6)
where {wti , t ∈ T } are measurement noise processes that are independent across time, independent of each other, and
independent of {wt0 , t ∈ T } and x1 .
The above evolution does not completely describe a dynamical control system, because we have not yet specified the data
available at each control station. In general, the random variable Iti available at control station i at time t will be a function
of all the past system variables {x[1,t] , y[1,t] , u[1,t−1] , w[1,t] }, i.e.,

Iti = ηti (x[1,t] , y[1,t] , u[1,t−1] , w[1,t] ), (10.7)

where we use the notation u = {u1 , . . . , uN } and x[1,t] = {x1 , . . . , xt }. The collection {Iti , i = 1, . . . , N , t ∈ T } is called
the information structure of the system, in analogy with Witsenhausen’s intrinsic model.
When T is finite, say equal to {1, . . . , T }, the above model is a special case of the sequential intrinsic model presented
above. The set {x1 , wt0 , wt1 , . . . , wtN , t ∈ T } denotes the primitive random variable with probability measure given by the
product measure of the marginal probabilities; the system has N × T DMs, one for each control station at each time. DM
(i, t) observes Iti and chooses uit . The information sub-fields J k are determined by {ηti , i = 1, . . . , N , t ∈ T }.
Some important information structures are
214 10 Decentralized and Multi-Agent Stochastic Control

1. Complete information sharing: In complete information sharing, each DM has access to present and past measure-
ments and past actions of all DMs. Such a system is equivalent to a centralized system.

Iti = {y[1,t] , u[1,t−1] }, t ∈ T.

2. Complete measurement sharing: In complete measurement sharing, each DM has access to the present and past
measurements of all DMs. Note that past control actions are not shared.

Iti = {y[1,t] }, t ∈ T.

3. Delayed information sharing In delayed information sharing, each DM has access to n-step delayed measurements
and control actions of all DMs.
(
i
{y[t−n+1,t] , ui[t−n+1,t−1] y[1,t−n] , u[1,t−n] }, t > n
Iti = i
(10.8)
{y[1,t] , ui[1,t−1] }, t≤n

4. Delayed measurement sharing In delayed measurement sharing, each DM has access to n-step delayed measure-
ments of all DMs. Note that control actions are not shared.
(
i
i {y[t−n+1,t] , ui[1,t−1] , y[1,t−n] }, t > n
It = i
{y[1,t] , ui[1,t−1] }, t≤n

5. Delayed control sharing In delayed control sharing, each DM has access to n-step delayed control actions of all
DMs. Note that measurements are not shared.
(
i
{y[1,t] , ui[t−n+1,t−1] , u[1,t−n] }, t > n
Iti = i
{y[1,t] , ui[1,t−1] }, t≤n

6. Periodic information sharing In periodic information sharing, the DMs share their measurements and control peri-
odically after every k time steps. No information is shared at other time instants.

i i
{y[⌊t/k⌋k,t] , u[⌊t/k⌋k,t] , y[1,⌊t/k⌋k] , u[1,⌊t/k⌋k] },

i
It = t≥k

 i i
{y[1,t] , u[1,t−1] }, t<k

7. Completely decentralized information In a completely decentralized system, no data is shared between the DMs.

Iti = {y[1,t]
i
, ui[1,t−1] }, t ∈ T.

In all the information structures given above, each DM has perfect recall (PR), that is, each DM has full memory of its
past information. In general, a DM need not have perfect recall. For example, a DM may only have access to its current
observation, in which case the information structure is

Iti = {yti }, t ∈ T. (10.9)

To complete the description of the team problem, we have to specify the loss function. For some applications, one may
have that the loss function is of additive form:
X
c(x[1,T ] , u[1,T ] ) := c(xt , ut ) (10.10)
t∈T
10.3 Solutions to Static Teams 215

where each term in the summation is known as the incremental (or stagewise) loss. The objective would be to choose
control policies γti such that uit = γti (Iti ) so as to minimize the expected loss (10.10).

10.3 Solutions to Static Teams

Definition 10.3.1 Given a static stochastic team problem {J; Γ i , i ∈ N }, a policy N -tuple γ ∈ Γ is stationary if (i) J(γ)
is finite, (ii) the N partial derivatives in the following equations are well defined, and (iii) γ satisfies these equations:

∇ui Eω|yi c(ω0 ; γ −i (y), ui ) |ui =γ i (yi ) = 0, a.s.


 
i ∈ N. (10.11)

There is a close connection between stationarity and person-by-person-optimality, as we discuss in the following. The re-
sults to be presented below are due to Krainak et. al. [195] and [349], generalizing Radner [260]. We follow the presentation
in [349], which also contains the proofs of the results.

Theorem 10.3.1 [260] [195] Let {J; Γ i , i ∈ N } be a static stochastic team problem where Ui = Rmi , i ∈ N , the loss
function c(ω0 , u) is convex and continuously differentiable in u a.s., and J(γ) is bounded from below on Γ. Let γ ∗ ∈ Γ
be a policy N -tuple with a finite cost (J(γ ∗ ) < ∞), and suppose that for every γ ∈ Γ such that J(γ) < ∞, the following
holds: X
E{∇ui c(ω0 ; u)|u=γ ∗ (y)[γ i (yi )−γ i∗ (yi )]}≥0,(10.12)
i∈N

where E{·} denotes the total expectation and the notation ∇ui c(ω0 ; γ ∗ (y)) means that the partial derivatives are evaluated
under policy γ ∗ . Then, γ ∗ is a team-optimal policy, and it is unique if c is strictly convex in u.

Proof Sketch. First, by the convexity of c, we obtain


1
c(ω0 ; γ ∗ (y) + α[γ(y) − γ ∗ (y)]) − c(ω0 ; γ ∗ (y)) ≤ c(ω0 ; γ(y)) − c(ω0 ; γ ∗ (y)),

α
for all α ∈ (0, 1]. Using the definition of J, this inequality can equivalently be written as (by taking the total expectation):
1
h(α) := [E{c(ω0 ; γ ∗ (y) + α[γ(y) − γ ∗ (y)])} − J(γ ∗ )] ≤ J(γ) − J(γ ∗ ),
α
where α ∈ (0, 1]. Note that both J(γ) and J(γ ∗ ) are finite, by hypothesis, and the first random variable (i.e., the first
loss function) also has a finite expectation for every α ∈ (0, 1] because of the bound provided by the inequality. Now,
due to the convexity of c, its finite integral, E{c(ω0 ; γ ∗ (y) + α[γ(y) − γ ∗ (y)])} is also convex in α. This leads to the
conclusion that (by a property of convex functionals that h(α) is a monotonically nonincreasing function as α ↓ 0, and
furthermore h(1) = J(γ) − J(γ ∗ ) is bounded (by hypothesis). It then follows from the monotone convergence theorem
that limα↓0 h(α) exists, and the limit and expectation operations can be interchanged. As a consequence of continuous
differentiability, this then leads to the inequality
N
X ∗
E{∇ui c(ω0 ; γ ∗ (y))[γ i (y i ) − γ i (y i )]} ≤ J(γ) − J(γ ∗ )
i=1

from which team-optimality of γ ∗ follows, since the left-hand-side is nonnegative, by (10.12).


If c were strictly convex in u, a.s., then all the inequalities above would be strict, for γ ̸= γ ∗ , thus leading to

J(γ ∗ ) < J(γ)

which implies that γ ∗ is the unique team-optimal solution. ⋄


Note that the conditions of Theorem 10.3.1 above do not include the stationarity of γ ∗ , and furthermore inequality (10.12)
may not generally be easy to check, since they involve all permissible policies γ (with finite cost).
216 10 Decentralized and Multi-Agent Stochastic Control

If the following N inequalities hold:

E{∇ui c(ω; γ ∗ (y))[γ i (y i ) − γ i∗ (y i )]} ≥ 0, i ∈ N, (10.13)

then (10.12) would also hold.


Then, either one of the following two conditions will achieve this objective [195] [349]:
(c.1) For all γ ∈ Γ such that J(γ) < ∞, the following random variables are integrable

∇ui c(ω0 ; γ ∗ (y))[γ i (y i ) − γ i∗ (y i )], i∈N

(c.2) Γ i is a Hilbert space for each i ∈ N , and J(γ) < ∞ for all γ ∈ Γ . Furthermore,

Eω|yi {∇ui c(ω0 ; γ ∗ (y)} ∈ Γ i , i ∈ N.

Here, (c.2) can directly be obtained from (c.1) if Γ i , i ∈ N , are taken as Hilbert spaces. Here we give it as a separate
condition because in some problems (such as linear quadratic—as we shall see shortly) (c.2) follows quite readily from the
problem formulations (due to the condition that a finite expected cost is attained under the considered policies).

Theorem 10.3.2 [195] [349] Let {J; Γ i , i ∈ N } be a static stochastic team problem which satisfies all the hypotheses of
Theorem 10.3.1, with the exception of the inequality (10.12). Instead of (10.12), let either (c.1) or (c.2) be satisfied. Then,
if γ ∗ ∈ Γ is a stationary policy it is also team optimal. Such a policy is unique if c(ω0 ; u) is strictly convex in u, a.s.

What needs to be shown is that under stationarity, (c.1) or (c.2) implies Theorem 10.3.1; this follows once again from the
law of the iterated expectations (Theorem 4.1.3); see [349]. If (c.1) holds, then for all i ∈ N ,
 
∗ i i i∗ i
E ∇ui c(ω0 ; γ (y))[γ (y ) − γ (y )]
  
∗ i i i∗ i i
= E E ∇ui c(ω0 ; γ (y))[γ (y ) − γ (y )] y
   
∗ i i i i∗ i
= E E ∇ui c(ω0 ; γ (y)) y (γ (y ) − γ (y ))

=0 (10.14)

under stationarity (where, again the order of expectation and differentiation is justified by the monotone convergence
theorem) and thus Theorem 10.3.2 holds.
To appreciate some of the fine points of Theorems 10.3.1 and 10.3.2, let us now consider the following example, which
was discussed by Radner (1962) [260], and Krainak et al. (1982) [195].

Example 10.2. Let N = 2, U1 = U2 = R, ξ = x be a Gaussian random variable with zero mean and unit variance
(∼ N (0, 1)), and the loss functional be given by
2
c(x; u1 , u2 ) = (u1 − u2 )2 ex + 2u1 u2 .

Note that c is strictly convex and continuously differentiable in (u1 , u2 ) for every value of x. Hence, if the true value of
x were known to both agents, the problem would admit a unique team optimal solution: u1 = u2 = 0, which is also
stationary. Since this team-optimal solution does not use the precise value of x, it is certainly optimal also under “no-
measurement” information at the decision makers. Note, however, that in this case the only pairs that make J(γ) finite, are
u1 = u2 = u ∈ R, since Z ∞
x2 1 x2
E[e ] = √ e+ 2 dx = ∞.
2π −∞
10.3 Solutions to Static Teams 217

The set of permissible policies not being an open set, we cannot talk about stationarity in this case. Theorem 10.3.1 (which
does not involve stationarity) is applicable here. Note also that for every u ∈ R, u1 = u2 = u is a pbp optimal solution,
but only one of these is team optimal.
Now, perhaps as a more practical case, consider the measurement scheme:

y 1 = x + w1 ; y 2 = x + w2

where w1 and w2 are independent random variables uniformly distributed on the interval [−1, 1], which are also indepen-
dent of x. Note that here the random state of nature, ξ, is chosen as (x, w1 , w2 )′ . Clearly, u1 = u2 = 0 is team-optimal for
this case also, but it is not obvious at the outset whether it is stationary or not. Toward this end, let us evaluate (10.11) for
i = 1 and with γ 2 (y 2 ) = 0:
2 2 2
(∂/∂u1 )Ex,y2 |y1 {(u1 )2 eξ } = (∂/∂u1 )[(u1 )2 Ex|y1 {eξ }] = 2u1 Ex|y1 {eξ }

where the last step follows because the conditional probability density of x given y 1 is nonzero only in a finite interval (thus
making the conditional expectation finite). By symmetry, it follows that both derivatives in (10.11) vanish at u1 = u2 = 0,
and hence the team-optimal solution is stationary. It is not difficult to see that in fact this is the only pair of stationary
policies. Note that all the hypotheses of Theorem 10.3.2 are satisfied here, under condition (c.2). ⋄

Quadratic-Gaussian teams

Given a probability space (Ω, F, PΩ ), and an associated vector-valued random variable ξ, let {J; Γ i , i ∈ N } be a static
stochastic team problem with the following specifications [349]:
(i) Ui = Rmi , i ∈ N ; i.e., the action spaces are unconstrained Euclidean spaces.
(ii) The loss function is a quadratic function of u for every ξ (where we use the notation L instead of c):
X ′ X ′
L(ξ; u) = ui Rij (ξ)uj + 2 ui ri (ξ) + c(ξ) (10.15)
i,j∈N i∈N


where Rij (ξ) is a matrix-valued random variable (with Rij = Rji ), ri (ξ) is a vector-valued random variable, and
c(ξ) is a random variable, all generated by measurable mappings on the random state of nature, ξ.
(iii) L(ξ; u) is strictly (and uniformly) convex in u a.s., i.e., there exists a positive scalar α such that, with R(ξ) defined
as a matrix comprised of N blocks, with the ij’th block given by Rij (ξ), the matrix R(ξ) − αI is positive definite
a.s., where I is the appropriate dimensional identity matrix.
(iv) R(ξ) is uniformly bounded above, i.e., there exists a positive scalar β such that the matrix βI − R(ξ) is positive
definite a.s.
(v) Y i = Rri , i ∈ N , i.e., the measurement spaces are Euclidean spaces.
(vi) y i = η i (ξ), i ∈ N , for some appropriate Borel measurable functions η i , i ∈ N .
(vii) Γ i is the (Hilbert) space of all Borel measurable mappings of γ i : Rri → Rmi , which have bounded second

moments, i.e., Eyi {γ i (y i )γ i (y i )} < ∞.
(viii) Eξ [ri′ (ξ)ri (ξ)] < ∞, i ∈ N ; Eξ [c(ξ)] < ∞.

Definition 10.3.2 A static stochastic team is quadratic if it satisfies (i)–(viii) above. It is a standard quadratic team if
furthermore the matrix R is constant for all ξ (i.e., it is deterministic). If, in addition, ξ is a Gaussian distributed random
vector, and ri (ξ) = Qi ξ, η i (ξ) = H i ξ, i ∈ N , for some deterministic matrices Qi , H i , i ∈ N , the decision problem is a
quadratic-Gaussian team (more widely known as a linear-quadratic-Gaussian (LQG) team under some further structure
on Qi and H i ). ⋄
218 10 Decentralized and Multi-Agent Stochastic Control

One class of quadratic teams for which the team-optimal solution can be obtained in closed form are those where the
random state of nature ξ is a Gaussian random vector. Let us decompose ξ into N + 1 block vectors
′ ′ ′
ξ = (x′ , y 1 , y 2 , . . . , y N )′ (10.16)

of dimensions r0 , r1 , r2 , . . . , rN , respectively. Being a Gaussian random vector, ξ is completely described in terms of its
mean value and covariance matrix, which we specify below:
′ ′
E[ξ] =: ξ¯ = (x̄′ , y 1 , . . . , y N ) (10.17)

cov (ξ) =: Σ, with [Σ]ij =: Σij , i, j = 0, 1, . . . , N (10.18)


[Σ]ij denotes the ij’th block of the matrix Σ of dimension ri × rj , which stands for the cross-variance between the i’th and
j’th block components of ξ. We further assume (in addition to the natural condition Σ ≥ 0) that Σii > 0 for i ∈ N , which
means that the measurement vectors y i ’s have nonsingular distributions. To complete the description of the quadratic-
Gaussian team, we finally take the linear terms ri (ξ) in the loss function (10.15) to be linear in x, which makes x the
“payoff relevant” part of the state of nature:
ri (ξ) = Di x, i ∈ N (10.19)
where Di is an (ri × r0 ) dimensional constant matrix.
In the characterization of the team-optimal solution for the quadratic-Gaussian team we will need the following important
result on the conditional distributions of Gaussian random vectors, generalizing our earlier results in Chapter 6.

Lemma 10.3.1 Let z and y be jointly Gaussian distributed random vectors with mean values z̄, ȳ, and covariance
 
Σzz Σzy
cov (z, y) = ′ ≥ 0, Σyy > 0. (10.20)
Σzy Σyy

Then, the conditional distribution of z given y is Gaussian, with mean


−1
E[z|y] = z̄ + Σzy Σyy (y − ȳ) (10.21)

and covariance
−1 ′
cov(z|y) = Σzz − Σzy Σyy Σzy (10.22)

The complete solution to the quadratic-Gaussian team is given in the following.

Theorem 10.3.3 [349] The quadratic-Gaussian team decision problem as formulated above admits a unique team-optimal
solution, that is affine in the measurement of each agent:

γ i∗ (y i ) = Π i (y i − y i ) + M i x̄, i ∈ N. (10.23)

Here, Π i is an (mi × ri ) matrix (i ∈ N ), uniquely solving the set of linear matrix equations:
X
Rii Π i Σii + Rij Π j Σji + Di Σ0i = 0, (10.24)
j∈N ,j̸=i

and M i is an (mi × r0 ) matrix for each i ∈ N , obtained as the unique solution of


X
Rij M j + Di = 0, i ∈ N . (10.25)
j∈N

Remark 10.3. The proof of this result follows immediately from Theorem 10.3.1 and noting that Condition (c.2) holds.
However, a Projection Theorem based concise proof can also be provided exploiting the quadratic nature of the problem
10.4 Static Reduction of Dynamic Teams: Policy-Dependent and Policy-Indepent Reductions 219

(see [349, p. 55], [261] and [138]), by defining the problem as an inner-product minimization and projection (onto the
closed subspace of decentralized control policies viewed as a product of individual policies of each DM) problem the
solution of which builds on an orthogonality condition.

An important application of the above result is the following static Linear Quadratic Gaussian Problem: Consider a two-
controller system evolving in Rn with the following description: Let x1 be Gaussian and x2 = Ax1 + B 1 u11 + B 2 u21 + w1

y11 = C 1 x1 + v11 ,

y12 = C 2 x1 + v12 ,
with w, v 1 , v 2 zero-mean, i.i.d. disturbances. For ρ1 , ρ2 > 0, let the goal be the minimization of
 
J(γ 1 , γ 2 ) = E ||x1 ||22 + ρ1 ||u11 ||22 + ρ2 ||u21 ||22 + ||x2 ||22 (10.26)

over the control policies of the form:


uit = µit (y1i ), i = 1, 2
For such a setting, optimal policies are linear.

10.4 Static Reduction of Dynamic Teams: Policy-Dependent and Policy-Indepent Reductions

Following Witsenhausen [331], we say that two information structures are equivalent if: (i) The policy spaces are equiv-
alent/isomorphic in the sense that policies under one information structure are realizable under the other information
structure, (ii) the costs achieved under equivalent policies are identical almost surely, and (iii) if there are constraints in the
admissible policies, the isomorphism among the policy spaces preserves the constraint conditions.
A large class of sequential team problems admit an equivalent information structure which is static. This is called the static
reduction of an information structure.

10.4.1 Static reduction I: Dynamic teams with quasi-classical information structures and their policy-dependent
static reduction

An important information structure which is not nonclassical, is of the quasi-classical type, also known as partially nested;
an IS is partially nested if an agent’s information at a particular stage t can depend on the action of some other agent
at some stage t′ ≤ t only if she also has access to the information of that agent at stage t′ . For such team problems with
partially nested information, one talks about precedence relationships among agents: an agent DM i is precedent to another
agent DM j (or DM i communicates to DM j), if the former agent’s actions affect the information of the latter, in which
case (to be partially nested) DM j has to have the information based on which the action-generating policy of DM i was
constructed.
For partially nested (or quasi-classical) information structures, static reduction has been studied by Ho and Chu in the
specific context of LQG systems [170] and for a class of non-linear systems satisfying restrictive invertibility properties
[171].
Under quasi-classical information, LQG stochastic team problems are tractable by conversion into equivalent static team
problems: Consider the following dynamic team with N agents, where each agent acts only once, with Ak, k ∈ N , having
the following measurement
X
yk = C k ξ + Dik ui , (10.27)
i:i→k

where ξ is an exogenous random variable picked by nature, and i → k denotes the precedence relation that the action of
Ai affects the information of Ak and ui is the action of Ai.
220 10 Decentralized and Multi-Agent Stochastic Control

If the information structure is quasi-classical, then

I k = {y k , {I i , i → k}}.

That is, Ak has access to the information available to all the signaling agents. Such an IS is equivalent to the IS I k = {ỹ k },
where ỹ k is a static measurement given by
 
k k i
ỹ = C ξ, {C ξ, i → k} . (10.28)

Such a conversion can be done provided that the policies adopted by the agents are deterministic, with the equivalence to
be interpreted in the sense that any deterministic policy measurable under the original IS being measurable also under the
new (static) IS and vice versa, since the actions are determined by the measurements. The restriction of using only deter-
ministic policies is, however, without any loss of optimality: with policies of all other agents fixed (possibly randomized)
no agent can benefit from randomized decisions in such team problems. We discussed this property of irrelevance of ran-
dom information/actions in optimal stochastic control in Chapter 5 in view of Blackwell’s Irrelevant Information Theorem
(see [346, Remark 2]).
This observation, made by Ho and Chu [170] leads to the following result.

Theorem 10.4.1 Consider an LQG system with a partially nested information structure. For such a system, optimal solu-
tions are affine (that is, linear plus a constant).

The linearity condition can be relaxed via the following more general condition, which is also due to Ho and Chu [171].

Assumption 10.4.1 Under a quasi-classical information structure, with

I k = {y k , {I i , i → k}}.

if y k = g(ξ, u[1,k−1] ), then the map g(·, u[1,k−1] ) : ξ 7→ y k is invertible.

Under this assumption, static reduction is possible.


Policy-dependence of the static reduction. In the above, under Assumption 10.4.1, while mapping the policies that are
equivalent under the dynamic setup to those that are expressed in terms of exogenous variables in the static-reduced form,
we note that the policies’ dependence on the exogenous variables explicitly depend on the policies adopted by the preceding
DMs. Accordingly, we refer to the static reduction of partially nested dynamic teams as policy-dependent static reduction
(as opposed to the policy-independent reduction to be presented in the following). Some restrictions and limitations due to
such policy-dependence will be studied later in the chapter.
If the control actions are also shared under the static measurement reductions, called static-measurements with control-
sharing reduction [277], even though this reduced information structure is not static in a strict sense, where only the mea-
surements are so, this reduction is [Link] follows since g(·, u[1,k−1] ) in Assumption 10.4.1 is invertible
given previous actions u[1,k−1] regardless of the policies of the previous decision makers.

Remark 10.4. Another class of dynamic team problems that can be converted into solvable dynamic optimization problems
are those where even though the information structure is nonclassical, there is no incentive for signaling because any
signaling from say agent Ai to agent Aj conveys information to the latter which is “cost irrelevant”, that is it does not lead
to any improvement in performance [342] [349].

10.4.2 Static reduction II: Non-classical information structures and their policy-independent reduction

In this sub-section, we introduce another static reduction method, due to Witsenhausen [331], applicable also to non-
classical information structures and call it a policy-independent static reduction, since this reduction does not depend on
the policies adopted.
10.4 Static Reduction of Dynamic Teams: Policy-Dependent and Policy-Indepent Reductions 221

For some of the results of the chapter, we need to go beyond a static reduction, and we will need to make the measurements
independent of each other as well as ω0 . This is not possible for every team which admits a static reduction, for example
quasi-classical team problems with LQG models [170] do not admit such a further reduction, since the measurements are
partially nested. Witsenhausen refers to such an information structure as independent static in [331, Section 4.2(e)].
Consider now dynamic team setting according to the intrinsic model where each DM t measures

y t = gt (ω0 , ωt , y 1 , . . . , y t−1 , u1 , . . . , ut−1 ),

and the decisions are generated by ut = γ t (y t ), with 1 ≤ t ≤ N . Here ω0 , ω1 , · · · , ωN are primitive (exogenous) variables.
We will indeed, for every 1 ≤ n ≤ N , view the relation

P (dy n |ω0 , y 1 , y 2 , · · · , y n−1 , u1 , u2 , · · · , un−1 ),

as a (controlled) stochastic kernel (to be defined later), and through standard stochastic realization results (see [143, Lemma
1.2] or [56, Lemma 3.1]), we can represent this kernel in a functional form through

y n = gn (ω0 , ωn , y 1 , y 2 , · · · , y n−1 , u1 , u2 , · · · , un−1 )

for some independent ωn and measurable gn .


This team admits an independent-measurements reduction provided that the following absolute continuity condition holds:
For every t ∈ N , there exists a function ft such that for all Borel S:

P (y t ∈ S|ω0 , u1 , u2 , · · · , ut−1 , y 1 , y 2 , · · · , y t−1 )


Z
= ft (y t , ω0 , u1 , u2 , · · · , ut−1 , y 1 , y 2 , · · · , y t−1 )Qt (dy t ), (10.29)
S

We can then write (since the action of each DM is determined by the measurement variables under a policy)

P (dω0 , dy, du)


YN  
= P (dω0 ) ft (y t , ω0 , u1 , u2 , · · · , ut−1 , y 1 , y 2 , · · · , y t−1 )Qt (dy t )1{γ t (yt )∈du} .
t=1

The cost function J(γ) can then be written as


Z N
Y
J(γ) = P (dω0 ) (ft (y t , ω0 , u1 , u2 , · · · , ut−1 , y 1 , y 2 , · · · , y t−1 )Qt (dy t ))c(ω0 , u), (10.30)
t=1

with uk = γ k (y k ) for 1 ≤ k ≤ N , and where now the measurement variables can be regarded as independent from each
other, and also from ω0 , and by incorporating the {ft } terms into c, we can obtain an equivalent static team problem.
Hence, the essential step is to appropriately adjust the probability space and the cost function.
The new cost function may now explicitly depend on the measurement values, such that
N
Y
cs (ω0 , y, u) = c(ω0 , u) ft (y t , ω0 , u1 , u2 , · · · , ut−1 , y 1 , y 2 , · · · , y t−1 ). (10.31)
t=1

Here we can reformulate even a static team to one which is, clearly still static, but now with independent measurements
which are also independent from the cost relevant exogenous variable ω0 .
Such a condition is in general not restrictive. Indeed, as Witsenhausen notes, a static reduction always holds when the
measurement variables take values from countableP set since a reference measure as in Qi above can be always constructed
−j
on the measurement space Y (e.g., Q (z) = j≥1 2 1{z=mj } where Yi = {mj , j ∈ N}) so that the absolute conti-
i i

nuity condition always holds. We refer the reader to [79] for relations with classical continuous-time stochastic control
222 10 Decentralized and Multi-Agent Stochastic Control

where the relation with Girsanov’s classical measure transformation [144] [33] is recognized. For discrete-time partially
observed stochastic control, similar arguments had been presented in Borkar [58], [62] again in the context of measure
transformation.

Remark 10.5. [Change of Measure Formula] Denote the joint probability measure on (ω0 , u1 , . . . , uN , y 1 , . . . , y N ) by
P , and the probability measure of ω0 by P0 . If the preceding absolute continuity condition (10.29) holds, then (under any
admissible policy profile γ 1 , · · · , γ N ) there exists a joint reference probability measure Q on (ω0 , u1 , . . . , uN , y 1 , . . . , y N )
such that
QN the probability measure P is absolutely continuous with respect to Q (P ≪ Q), so that for every Borel set A in
(Ω0 × i=1 (Ui × Yi ))

dP
Z
P (A) = Q(dω0 , du1 , . . . , duN , dy 1 , . . . , dy N ), (10.32)
A dQ

where the reference probability measure


N
Y
Q(dω0 , du1 , . . . , duN , dy 1 , . . . , dy N ) := P0 (dω0 ) Qi (dy i )1{γ i (yi )∈dui } , (10.33)
i=1

leads to a Radon-Nikodym derivative, which is policy-independent:


N
dP Y
(ω0 , u1 , . . . , u1 , y 1 . . . , y N ) = f i (y i , ω0 , u1 , . . . , ui−1 , y 1 , . . . , y i−1 ). (10.34)
dQ i=1

Indeed, one may slightly relax the condition in (10.29) (which requires the absolute continuity to hold for all ω0 , u1 , · · · , ut−1 , y 1 , · · · , y t−1 ),
to an almost sure existence condition of a derivative under a reference measure, in the sense that (10.34) holds.

Witsenhausen’s Counterexample and its static reduction

The celebrated Witsenhausen’s counterexample [325] is a dynamic non-classical team problem


Suppose x and w1 are two independent, zero-mean Gaussian random variables with variance σ 2 and 1 so that

y 0 = x, y 1 = u0 + w1

u0 = γ0 (x), u1 = γ1 (y).
with the performance criterion:
QW (x, u0 , u1 ) = k(u0 − x)2 + (u1 − u0 )2 , (10.35)

This can also be viewed as a standard discrete-time two-stage stochastic optimal control problem, with state equations (see
Figure 10.3)

x 1 = x 0 + v0 , x2 = x1 − v1 , (10.36)

measurement equations

w1
x u0 y u1
γ0 γ1

Fig. 10.2: Flow of information in Witsenhausen’s counterexample.


10.4 Static Reduction of Dynamic Teams: Policy-Dependent and Policy-Indepent Reductions 223

y0 = x0 , y1 = x1 + w1 , (10.37)

and memoryless controls

v0 = γ0 (y0 ) , v1 = γ1 (y1 ) , (10.38)

where µ0 and µ1 are the instantaneous measurement output control policies at stages 0 and 1, respectively. This becomes
equivalent to the earlier formulation in view of the correspondences

u 0 = x 0 + v0 , u1 = v1 , x = x0 , w = w1 , y = y1 ,

if we pick the cost function as

Q̃(x2 , v0 ) = (x2 )2 + k(v0 )2 ≡ QW (x1 − v0 , x1 , x1 − x2 ) .

This problem is described by a linear system; all primitive variables are Gaussian and the performance criterion is quadratic,
yet linear policies are not optimal. We note that this is a non-convex problem [351] and thus variational methods do not
necessarily lead to optimality. In fact, we don’t even have a good lower bound on the optimal cost for Witsenhausen’s
counterexample even though approximation results exist (see [274] for a detailed discussion).
The static reduction for Witsenhausen’s counterexample proceeds as follows:
Z
(k(u1 − y 1 )2 + (u1 − u2 )2 )Q(dy 1 )γ 1 (du1 |y 1 )γ1 (du2 |y 2 )P (dy 2 |u1 )
Z
= (k(u1 − y 1 )2 + (u1 − u2 )2 )Q(dy 1 )γ 1 (du1 |y 1 )γ1 (du2 |y 2 )η(y 2 − u1 )dy 2
Z  2 1 2

2 2 1 1 1 2 2 η(y − u )dy
= (ku0 + (u0 − u1 ) )γ (du |y )γ1 (du |y ) Q(dy 1 )η(y 2 )dy 2
η(y 2 )
Z  2 1 2

2 2 1 1 1 2 2 η(y − u )dy
= (ku0 + (u0 − u1 ) )γ (du |y )γ1 (du |y ) Q(dy 1 )Q(dy 2 ) (10.39)
η(y 2 )

where Q denotes a Gaussian measure with zero mean and unit variance and η its density.

10.4.3 Equivalent static reductions preserve optimality but may not person-by-person-optimality or stationarity

We showed earlier that static reduction can be a very useful method for deriving and studying properties of optimal poli-
cies in stochastic teams. We note, however, that static reduction has its limitations when the solution concept is not global
optimality but only person-by-person-optimality or stationarity. This, in particular, is a concern for policy-dependent re-
ductions.

Theorem 10.4.2 [277] Consider a stochastic dynamic team with a policy-independent static reduction.
(i) A policy γ ∗ is pbp optimal (globally optimal) for a dynamic team if and only if γ ∗ is pbp optimal (globally optimal)
for a policy-independent static reduction of the dynamic team;

w1

y0 v0 y1 v1
γ0 γ1

x0 x1 x2

Fig. 10.3: Witsenhausen’s counterexample in two-stage state-space linear stochastic control form.
224 10 Decentralized and Multi-Agent Stochastic Control

(ii) Let a policy γ ∗ satisfy P -almost surely


 
−i∗
∇ui EQγ dP
dQ yi =0 ∀i ∈ N , (10.40)
ui =γ i∗ (y i )

∗ ∗
where dP
dQ is defined in (10.34). Then, γ is stationary for dynamic team if and only if γ is stationary for a policy-
independent static reduction of dynamic team.

Theorem 10.4.3 [277] Consider a stochastic dynamic team with partially nested information structure. Let Assumption
10.4.1 hold. Then:
(i) γ D,∗ is a globally optimal policy for for the dynamic team if and only if its static equivalent γ S,∗ is a globally
optimal policy for its static reduction, under the policy-dependent static reduction.
(ii) If γ D,∗ is a stationary (pbp optimal) policy for (P D ), then its static equivalent γ S,∗ is not necessarily a stationary
(pbp optimal) policy for (P S ) under the policy-dependent static reduction;
(iii)If γ S,∗ is a stationary (pbp optimal) policy for a static reduced dynamic team, then γ D,∗ , satisfying the policy-
dependent static reduction relation, is not necessarily pbp optimal for the dynamic team problem.

10.4.4 All stochastic dynamic teams are nearly static (with independent measurements) reducible

Now that we have seen the benefits of static reduction, a natural question arises as to whether we can perturb any stochastic
dynamic team by adding some additive noise to the measurements to make it static-reducible with arbitrarily small error
in the optimal cost: That is, are all dynamic team problems ϵ-away from being static reducible as far as optimal cost is
concerned for any ϵ > 0? This is indeed the case, see [172].

10.5 Expansion of information Structures: A recipe for identifying sufficient information

We start with a general result on optimum-performance equivalence of two stochastic dynamic teams with different infor-
mation structures. This is in fact a result which has a very simple proof, but it is quite effective as we will see shortly.

Proposition 10.5.1 Let D1 and D2 be two stochastic dynamic teams with the same loss function, and differing only in
their information structures, η 1 and η 2 , respectively, with corresponding composite strategy spaces Γ1 and Γ2 , such that
Γ2 ⊆ Γ1 . Let D1 admit a team-optimal solution, denoted by γ ∗1 ∈ Γ1 , with the further property that γ ∗1 ∈ Γ2 . Then γ ∗1
also solves D2 .

A recipe for utilizing the result above would be [349]:


Given a team problem, say D2 , with IS η 2 , which is presumably difficult to solve, obtain a finer IS η 1 , and solve
the team problem under this expanded IS (assuming that this new team problem is easier to solve). Then, if the
team-optimal solution here is adapted to the sigma-field generated by the original coarser IS, it solves also the
original problem D2 .

10.6 Convexity of Decentralized Stochastic Control Problems

We have already seen the utility of convexity in team theoretic problems earlier in the chapter, e.g. in Theorem 10.3.1. We
will study convexity further in this section.
10.6 Convexity of Decentralized Stochastic Control Problems 225

10.6.1 Convexity of static team problems and an equivalent representation of cost functions

Definition 10.6.1 A (static or dynamic) team problem is convex on Γ if J(γ) < ∞ for all γ ∈ Γ and for any α ∈
(0, 1), γ 1 , γ 2 ∈ Γ:
J(αγ 1 + (1 − α)γ 2 ) ≤ αJ(γ 1 ) + (1 − α)J(γ 2 )

We state the following immediate result without proof, more general refinements will be stated later in the chapter.

Theorem 10.6.1 Consider a static team. J(γ) is convex if c(ω0 , u) is convex in u for all ω0 , provided that J(γ) < ∞ for
all γ ∈ Γ.

The condition in Theorem 10.6.1 is not tight, however, due to information structure and measurability aspects.

Example 10.6. Consider Ω = [0, 1] and let P be the uniform distribution on Ω, with N = 2, U1 = U2 = [1, 2]. Let:
  p p 
c(ω, u1 , u2 ) = 1{ω∈[0,0.9]} (u1 − 2)2 + (u2 − 2)2 + 1{ω∈(0.9,1]} 1 + u1 + 1 + u2

Now, suppose further that I 1 = I 2 = η 1 (ω) = η 2 (ω) = 1{ω∈[0,0.1)} . It follows that here the team problem is convex,
even though c(ω, u1 , u2 ) is not convex on {ω : ω ∈ (0.9, 1]}, which has a non-zero probability measure. To see this, note
that one may view this optimization problem as J(u11 , u12 ; u21 , u22 ) where uij = γ i (ωj ), with ω1 ≡ {ω : ω ∈ [0, 0.1)} and
ω2 ≡ {ω : ω ∈ [0.1, 1]}. It follows that
X q
J(u11 , u12 ; u21 , u22 ) = 0.1(ui1 − 2)2 + 0.8(ui2 − 2)2 + 0.1( ui2 + 1)
i=1,2

The Hessian of J is a diagonal matrix with strictly positive entries, leading to the convexity of the problem.

In the following, we will make use of the fact that uk ↔ y k ↔ {y−k , ω} form a Markov chain almost surely. Before
proceeding further, let us note that the join of two σ-fields over some set X is the coarsest σ-field containing both. The
meet ofTtwo σ-fields is the finest σ-field which is a subset of both. Let F i be the σ-field generated by η i over Ω, and let
Fc = k F k be the meet of these fields, this is termed as common knowledge by Aumann [22] for finite probabilities
spaces. In addition, let Fj be the join of the σ-field, denoted with Fj = k F k .
S

In the following, as earlier in the chapter, we assume that the measurement and the control action spaces are standard Borel.
An equivalent representation of the cost through iterated expectations. Let us express the expected cost under a given
measurable team policy γ as follows. With the interpretation that P (uk ∈ ·|y k ) = 1{uk =γ k (yk )∈·} , we obtain from the law
of the iterated expectations that
 
E[c(ω0 , u)] = E E[c(ω0 , u)|y] (10.41)

Under any measurable policy, given y, u is specified. Thus, with

c̃(y 1 , · · · , y N , u1 , · · · , uN ) := E[c(ω0 , u)|y],


 
1 1
N N
P (dω0 |y)c(ω0 , u , · · · , u ) , the cost function becomes E[c̃(y 1 , · · · , y N , u1 , · · · , uN )].
1 N
R
that is c̃(y , · · · , y , u , · · · , u ) =
We will use this representation in the following.

Theorem 10.6.2 (i)If a team problem is convex, then

E[c(ω0 , u)|Fc ]
226 10 Decentralized and Multi-Agent Stochastic Control

is convex in u almost surely.


(ii) If
E[c(ω0 , u)|Fj ]
is convex in u almost surely, then the team problem is convex on the set of team policies that satisfy J(γ) < ∞.

Proof. (i) We will show the contra-positive. Let B be a Borel set such that P (B) > 0, B ∈ Fc , and E[c(ω0 , u)|B] be
non-convex so that there exist u and u′ and λ ∈ (0, 1) such that

E[c(ω0 , λu + (1 − λ)u′ )|B] > λE[c(ω0 , u)|B] + (1 − λ)E[c(ω0 , u′ )|B]

Now, let γ and γ be two team policies so that these only differ on B; and on B γ = u and γ = u′ . Such measurable
policies exist, for example by taking γ(ω) = {0, 0, · · · , 0} when ω ∈
/ B. These policies are both Borel measurable and are
admissible given the information structure. Then

J(λγ + (1 − λ)γ) > λJ(γ) + (1 − λ)J(γ)

and convexity fails.

(ii) We adopt the equivalent representation (10.41) in this part of the proof. Note that under any measurable policy, the
random variable c̃(y 1 , · · · , y N , u1 , · · · , uN ) is measurable on the σ-field generated by y and thus the join σ-field. The
proof then follows from the following. Consider two policies γ and γ̄ with finite expected costs. It follows then that

J(λγ + (1 − λ)γ̄)
Z
= P (dy)c̃(y 1 , · · · , y N , λγ 1 (y 1 ) + (1 − λ)γ̄ 1 (y 1 ), · · · , λγ N (y N ) + (1 − λ)γ̄ N (y N ))
Z 
≤ P (dy) λc̃(y 1 , · · · , y N , γ 1 (y 1 ), · · · , γ N (y N ))

+(1 − λ)c̃(y 1 , · · · , y N , γ̄ 1 (y 1 ), · · · , γ̄ N (y N ))

= λJ(γ) + (1 − λ)J(γ̄)


It can be observed that Example 10.6 satisfies the conditions of Theorem 10.6.2. These conditions will also be used to
study Witsenhausen’s counterexample [325] later in the chapter.

A generalization of Radner and Krainak et. al.’s theorems

We provide a generalization of Radner’s or Krainak et al.’s theorem by utilizing an information structure dependent nature
of convexity. For example, Radner or Krainak [Link]’s theorems are not applicable to Example 10.6.

Theorem 10.6.3 Let {J; Γ i , i ∈ N } be a static stochastic team problem, the loss function E[c(ω0 , u)|Fj ] is convex and
continuously differentiable in u almost surely. Let γ ∗ ∈ Γ be a policy N -tuple with a finite cost (J(γ ∗ ) < ∞), and suppose
that for every γ ∈ Γ such that J(γ) < ∞, the following holds:
X
E{∇ui c̃(y, γ ∗ (y))[γ i (y i ) − γ i∗ (y i )]} ≥ 0, (10.42)
i∈N

where c̃(y, u) := E[c(ω0 , u)|Fj ]. Then, γ ∗ is a team-optimal policy, and it is unique if c̃(y, u) is strictly convex in u
almost surely.
10.6 Convexity of Decentralized Stochastic Control Problems 227

Proof. The proof follows by defining the new loss function:

c̃(y, u) = E[c(ω0 , u)|Fj ],

almost surely. The result then follows as in Theorem 10.3.1. ⋄

Theorem 10.6.4 Let {J; Γ i , i ∈ N } be a static stochastic team problem which satisfies all the hypotheses of Theo-
rem 10.6.3, with the exception of inequality (10.42). Instead of (10.42), let either (c.5) or (c.6) be satisfied with c replaced
with c̃. Then, if γ ∗ ∈ Γ is a stationary policy it is also team optimal. Such a policy is unique if E[c(ω0 , u)|Fj ] is strictly
convex in u, a.s.

Proof. The proof follows by defining the new loss function c̃ as in the proof of Theorem 10.6.3, and following Theorem
10.3.2. ⋄

10.6.2 Convexity of Sequential Dynamic Teams

Convexity of the reduced model

The static reduction of a sequential dynamic team problem, if exists, is not unique. However, the following holds: Either all
of the static reductions are convex or none is. This holds under a minor technicality for quasi-classical patterns. Here, first
the information is to be expanded to allow for control sharing. Thus, we can state that a stochastic dynamic team problem
with a static reduction is convex if and only if its static reduction is.
Non-convexity of the Witsenhausen counterexample and its variants. Consider the celebrated Witsenhausen’s coun-
terexample [325]: This is a dynamic non-classical team problem with y 1 and w1 zero-mean independent Gaussian
random variables with unit variance and u1 = γ 1 (y 1 ), u2 = γ 2 (u1 + w1 ) and the cost function c(ω, u1 , u2 ) =
k 2 (y 1 − u1 )2 + (u1 − u2 )2 for some k > 0: The static reduction is given in (10.39).
Another interesting example is the point-to-point communication problem: Here, the setup is exactly as in the Witsen-
hausen’s counterexample, but c(ω, u1 , u2 ) = k 2 (u1 )2 + (y 1 − u2 )2 . This problem is a peculiar one in that, even though the
information structure is non-classical, and is non-convex; an optimal encoder and decoder is linear. A proof of this result
builds on information theoretic ideas, such as the data-processing inequality (see Chapters 3, 11 in [349] for a detailed
account). In this case, the reduction (10.30) writes as:
Z
(k(u1 )2 + (y 1 − u2 )2 )Q(dy 1 )γ 1 (du1 |y 1 )γ1 (du2 |y 2 )P (dy 2 |u1 )
Z  2 1 2

1 2 1 2 2 1 1 1 1 2 2 η(y − u )dy
= (k(u ) + (y − u ) )Q(dy )γ (du |y )γ1 (du |y ) Q(dy 1 )Q(dy 2 )
η(y 2 )
(10.43)

Consider the static reduction of Witsenhausen’s counterexample and the Gaussian signaling problem (10.39)-(10.43). For
2
both (10.39) and (10.43), using the fact that e−x is not a convex function, we recognize that this problem is not convex by
Theorem 10.6.2(i) (with the common knowledge/information being the trivial σ-algebra consisting of the empty set and its
complement).
We note that Witsenhausen states without proof in [325] (p. 134) that the counterexample is non-convex in γ 1 for every
optimal γ 2 (selected as a best response to γ 1 ). The discussion above can be viewed as an explicit proof for this result. Note
also that for both problems above, linear policies contain person by person optimal policies, but this does not imply global
optimality. For the first problem, Witsenhausen had shown the suboptimality of linear policies. For the second problem
(10.43), however, linear policies are indeed optimal.
228 10 Decentralized and Multi-Agent Stochastic Control

Partially nested information structures: Convexity of the reduced model

As reviewed earlier, an important information structure which is not nonclassical, is of the partially nested type. For such
team problems with partially nested information, a static reduction exists under certain invertibility conditions as discussed
earlier. For such problems, the cost function is not altered by the static reduction. This leads to the following result.

Theorem 10.6.5 Consider a partially nested stochastic dynamic team which admits a static reduction where the cost func-
tion c(ω0 , u) convex in u. If the information structure is expanded to also include control sharing whenever measurements
are shared under the partially nested information structure, then the team problem is convex.

See [351]. We note that Ho and Chu [170] established this result that for the special setup involving the partially nested
LQG teams. In this case, optimal policies are linear through an equivalence to static teams.

10.6.3 Symmetric Team Problems: Optimality of Symmetric Policies

If a team problem, static or dynamic, is convex and symmetric (i.e., exchangeable; meaning that any permutation of DM
policies does not alter the induced expected cost), then an optimal team policy can be taken to be symmetric across agents
without loss.
In the absence of convexity, one can only show exchangeability of an optimal team policy [278–280].

10.7 The Strategic Measures Approach to Decentralized Stochastic Control

For classical stochastic control problems, strategic measures were defined (see [284], [253], [116] and [123]) as the set
of probability measures induced on the product (sequence) spaces of the states, measurements, and actions; that is, given
an initial state distribution and a policy, one can uniquely define a probability measure on the product space of the states,
measurements, and actions. Certain measurability, compactness, and convexity properties of strategic measures for classical
stochastic control problems were studied in [49, 116, 123, 253].
In [351], strategic measures for decentralized stochastic control problems were introduced and many of their properties
were established. For decentralized stochastic control problems, considering the set of strategic measures along with com-
pactification and/or convexification of these sets of measures through introducing private and/or common randomness al-
low one to place operationally flexible topologies (such as those leading to a standard Borel space, e.g., weak convergence
topology, among others) on the set of strategic measures, as we will study in the following.

10.7.1 Measurable policies as a subset of randomized policies and strategic measures

A common method in control theory is to view a measurable policy as a special case of relaxed policies where relaxation
is often employed by randomization. Such an approach has been ubiquitously adopted in various fields often with different
terminology (e.g., relaxed controls (Young topology) in optimal deterministic control [223] [340], distributional strategies
in economics [237] [229], local hidden variables in quantum information theory etc.)
We recall here the following representation result [56]. Let X, M be Borel spaces. Let the notation P(X) denote the set of
probability measures on X. Consider the set of probability measures

Θ := ζ ∈ P(X × M) : ζ(dx, dm) = P (dx) Qf (dm|x), Qf (·|x) = 1{f (x)∈·} , f : X → M ,




on X × M having fixed input marginal P on X and the stochastic kernel from X to M is realized by some measurable
function f : X → M. We equip this set with weak convergence topology. This set is the (Borel measurable) set of the
extreme points of the set of probability measures on X × M with a fixed marginal P on X. For compact M, the Borel
10.7 The Strategic Measures Approach to Decentralized Stochastic Control 229

measurability of Θ follows from [251] since the set of probability measures on X × M with a fixed marginal P on X
is a convex and compact set in a complete separable metric space, and therefore, the set of its extreme points is Borel
measurable. Moreover, the non-compact case holds by [56, Lemma 2.3]. Furthermore, given a fixed marginal P on X, any
stochastic kernel Q from X to M can be identified by a probability measure ξ ∈ P(Θ) such that
Z
Q(·|x) = ξ(dQf ) Qf (·|x). (10.44)
Θ

In particular, a stochastic kernel can thus be viewed as an integral representation over probability measures induced by
deterministic policies.
For a team setup, for any DM k, let

k
Θk := ζ ∈ P(Yk × Uk ) : ζ = Pk Qγ ,

k
Qγ (·|y k ) = 1{γ k (yk )∈·} , γ k ∈ Γ k , Pk (·) = P (y k ∈ ·) .

For a static team, Pk would be fixed; that is, independent of the policies of the preceding DMs. Therefore, in static case, in
view of (E.4), any element ζ ∈ P(Yk × Uk ) with fixed marginal Pk on Yk can be expressed as the mixture of Θk
Z
ζ(A) = ξ k (dQ) Q(A), A ∈ B(Yk × Uk ), (10.45)
Θk

for some ξ ∈ P(Θk ). In the sequel, we generalize this idea to the set of strategic measures induced by measurable policies
and define various relaxed policies that are obtained as a mixture of measurable policies. Indeed, instead of viewing N -
tuple of policies as the joint strategy of DMs, we regard the induced probability distribution on the product space of state,
measurements, and actions as the joint strategy and name it strategic measure.

10.7.2 Sets of strategic measures for static teams

QN
Consider a static team problem defined under Witsenhausen’s intrinsic model. In the following, B = B 0 × k=1 (A
k
×B k )
QN
are used to denote the cylindrical Borel sets in Ω0 × k=1 (Yk × Uk ).

Let LA (µ) be the  induced by all admissible measurable policies with (ω0 , y) ∼ µ; that is,
 set of strategic measures
QN
P ∈ LA (µ) ⊂ P Ω0 × k=1 (Yk × Uk ) if and only if

Z N
Y
P (B) = QN µ(dω0 , dy) 1{uk =γ k (yk )∈B k } , (10.46)
B0 × Ak k=1
k=1

 QN 
for all cylindrical B ∈ B Ω0 × k=1 (Yk × Uk ) , where γ k ∈ Γ k for k = 1, . . . , N . Let LA (µ, γ) be the strategic
measure under a particular strategy γ ∈ Γ.
The first relaxation is obtained via individual randomization of policies. Namely, let LR (µ) be the set of strategic measures
induced by all individually randomized team policies where ω0 , y ∼ µ; that is,
  N
Y  Z N
Y 
k k k k k
LR (µ) := P ∈ P Ω0 × (Y × U ) : P (B) = µ(dω0 , dy) Π (du |y ) ,
k=1 B k=1

where Π k takes place from the set of stochastic kernels from Yk to Uk for each k = 1, . . . , N .
Another relaxation, which is stronger than the former one, is obtained by taking the mixture of the elements of LA (µ). To
this end, define Υ = [0, 1]N . We then let
230 10 Decentralized and Multi-Agent Stochastic Control
  N
Y  Z 
LC (µ) := P ∈ P Ω0 × (Yk × Uk ) : P (B) = η(dz)LA (µ, γ(z))(B), η ∈ P(Υ ) ,
k=1

where γ(z) denotes a collection of team policies measurably parametrized by z ∈ Υ so that the map LA (µ, γ(·)) : Υ →
 
QN k k
LA (µ) is Borel measurable as LA (µ) is a Borel subset of P Ω0 × k=1 (Y × U ) under weak convergence topology
(as we will see in Theorem 10.10).
Let LCR denote the set of strategic measures that are induced by some fixed but common independent randomness and
arbitrary private independent randomness; that is,
  N
Y 
k k
LCR (µ) := P ∈ P Ω0 × (Y × U ) :
k=1
Z Y 
k k k
P (B) = η(dz)µ(dω0 , dy) Π (du |y , z) ,
B×Υ k

where Π k takes place from the set of stochastic kernels from Yk × Υ to Uk for each k = 1, . . . , N . Here, the common
randomness η is fixed.
Let LCCR denote the set of strategic measures that are induced by some arbitrary but common independent randomness
and arbitrary private independent randomness, as in LC (µ); that is,
  N
Y 
LCCR (µ) := P ∈ P Ω0 × (Yk × Uk ) :
k=1
Z Y 
P (B) = η(dz)µ(dω0 , dy) Π k (duk |y k , z), η ∈ P(Υ ) ,
B×Υ k

where Π k takes place from the set of stochastic kernels from Yk × Υ to Uk for each k = 1, . . . , N . Here, the common
randomness η is arbitrary, unlike LCR (µ). The following result, essentially from [351], states some structural results about
above-defined sets of strategic measures. In particular, it establishes convexity related properties of these sets.
There also exist further convex relaxations: Quantum Relaxations, Non-Signaling Relaxations and Local-Markov Relax-
ations. We do not discuss these in these notes.

Theorem 10.7. Consider a static team problem. Then, we have the following characterizations.
(i) LR (µ) has the following representation:
  N
Y  Z
k k
LR (µ) = P ∈ P Ω0 × (Y × U ) : P (B) = U (dz)LA (µ, γ(z))(B),
k=1
Y 
U ∈ P(Υ ), U (dv1 , · · · , dvN ) = ηk (dvk ), ηk ∈ P([0, 1]) ;
s

that is, U ∈ P(Υ ) is constructed by the product of N independent random variables on [0, 1].
(ii) LC (µ) = LCCR (µ) and this is a convex set. The set of extreme points of LC (µ) is LA (µ). Furthermore, LR (µ) ⊂
LC (µ).
(iii) We have the following equalities:
Z Z Z
inf J(γ) = inf P (ds)c(s) = inf P (ds)c(s) = inf P (ds)c(s).
γ∈Γ P ∈LA (µ) P ∈LR (µ) P ∈LC (µ)

In particular, deterministic policies are optimal among the randomized class. In other words, individual and common
randomness does not improve the optimal team cost.
10.7 The Strategic Measures Approach to Decentralized Stochastic Control 231

(iv) The sets LR (µ) and LCR (µ) are not convex. In particular, the presence of independent or (fixed) common randomness
does not convexify the set of strategic measures.
(v) LR (µ) and LC (µ) are not necessarily weakly closed.

10.7.3 Sets of strategic measures for dynamic teams in the absence of static reduction

Note that if the dynamic team setup admits a static reduction (in particular independent static reduction), then one can
define strategic measures by considering equivalent static problem and characterize the convexity properties of the set of
strategic measures, as done in the previous section. In this section, we suppose that dynamic team does not admit a static
reduction. Let µ be the distribution of ω0 . Recall that in dynamic setup, the distribution of measurements y is not fixed
as opposed to the static case. In this case, we present the following characterization for strategic measures in dynamic
sequential teams. Let, for all n ∈ N ,

hn = {ω0 , y 1 , u1 , · · · , y n−1 , un−1 , y n , un },

and pn (dy n |hn−1 ) := P (dy n |hn−1 ) be the transition kernel characterizing the measurements of DM n according to the
intrinsic model. We note that this may be obtained by the relation:

pn (y n ∈ ·| ω0 , y 1 , u1 , · · · , y n−1 , un−1 )
 
:= P η i (ω, u[1,i−1] ) ∈ · ω0 , y 1 , u1 , · · · , y n−1 , un−1
 
= P g n (ω0 , ωn , u1 , · · · , un−1 ) ∈ · ω0 , y 1 , u1 , · · · , y n−1 , un−1 . (10.47)

Note that once a policy is fixed, pn (dy n |hn−1 ) represents the conditional distribution of y n given the past history hn−1 .
Let LA (µ) be the set of strategic measures induced by measurable policies and let LR (µ) be the set of strategic measures
induced by individually randomized policies for the dynamic team. We have the following characterizations of LA (µ) and
LR (µ) that are quite useful when establishing the closedness of these sets.

Theorem 10.8 ( [351, Theorem 2.2]). Consider a dynamic team problem that does not admit a static reduction. Then, we
have the following characterizations.
 
QN
(i) A probability measure P ∈ P Ω0 × k=1 (Yk × Uk ) is a strategic measure induced by a measurable policy (that is
in LA (µ)) if and only if, for every n = 1, . . . , N , we have
Z Z Z 
n n
P (dhn−1 , dy ) g(hn−1 , y ) = P (dhn−1 ) g(hn−1 , z) pn (dz|hn−1 )
Yn

and Z Z Z 
P (dhn ) g(hn−1 , y n , un ) = P (dhn−1 , dy n ) g(hn−1 , y n , a) 1{γ n (yn )∈da} ,
Un

for all continuous and bounded function


 g with appropriatearguments, where P (dω0 ) = µ(dw0 ) and γ n ∈ Γ n .
QN
(ii) A probability measure P ∈ P Ω0 × k=1 (Yk × Uk ) is a strategic measure induced by a individually randomized
policy (that is in LR (µ)) if and only if, for every n = 1, . . . , N , we have
Z Z Z 
P (dhn−1 , dy n ) g(hn−1 , y n ) = P (dhn−1 ) g(hn−1 , z) pn (dz|hn−1 ) (10.48)
Yn

and
Z Z Z 
P (dhn ) g(hn−1 , y n , un ) = P (dhn−1 , dy n ) g(hn−1 , y n , an ) Π n (dan |y n ) (10.49)
Un
232 10 Decentralized and Multi-Agent Stochastic Control

for all continuous and bounded function g with appropriate arguments, where P (dω0 ) = µ(dw0 ) and Π n is a stochastic
kernel from Yn to Un .

Remark 10.9. A result similar to Theorem 10.7 can also be stated for the dynamic case, in particular with regard to LA (µ)
being the set of extreme points of the convex hull of LR (µ). The reader is referred to [351, Theorem 2.3] which essentially
establishes this; see also [122, Theorem 1.c] for related discussions.
A celebrated result in economics theory, known as Kuhn’s theorem [198], notes that the convex hull of admissible (i.e.
those in LA (µ)) strategic measures (hence LC (µ)) is equivalent to LR (µ) when the information structure is classical. We
can thus state that this does not apply in the absence of classical-ness, as LR (µ) would not be convex (if the information
structure is not classical, then convexity fails [351, p.12]), but the convex hull of admissible policies is, by definition,
convex; but the convex hull of LR (µ) is LC (µ).

10.7.4 Measurability properties of sets of strategic measures


As noted earlier, the set LA (µ) is a Borel subset of P Ω0 × k (Yk × Uk ) under weak convergence topology. The same
Q
is true for LR (µ), which is stated in the following theorem. This result will be crucial in the analysis to follow.

Theorem 10.10 ( [351, Theorem 2.10]). Consider a sequential (static or dynamic) team.
(i) The set of strategic measures LR (µ) is Borel when viewed as a subset of the space of probability measures on
QN
Ω0 × k=1 (Yk × Uk ) under the topology of weak convergence.
(ii) The set of strategic measures LA (µ) is Borel when viewed as a subset of the space of probability measures on
QN
Ω0 × k=1 (Yk × Uk ) under the topology of weak convergence.

For further properties of the sets of strategic measures, see [351].

10.8 Existence of Optimal Solutions

The following theorem states a general existence result for static teams and for dynamic teams admitting static reduction.
Its proof depends on Weierstrass Extreme Value Theorem.

Theorem 10.11. Consider a static team or the static reduction of a dynamic team with c denoting the cost function. Let
c be lower semi-continuous in u for every fixed ω0 , y and LR (µ) or LC (µ) be a compact set under weak convergence
topology. Then, there exists an optimal team policy. This policy can be chosen deterministic and hence induces a strategic
measure in LA (µ).

Remark 10.12. Since the cost function cs in independent static reduction of a dynamic team also depends on the measure-
ments y, we include y as an argument to the cost function c in the previous theorem.

Theorem 10.13. [346, Theorem 5.2] Consider a static or a dynamic team that admits an independent static reduction.
Let c be lower semi-continuous in u for any ω0 , y. Suppose further that Ui is σ−compact (that is, Ui = ∪n Kn for a
countable collection of increasing compact sets Kn ) and, without any loss, the control laws can be restricted to those with
E[ϕi (ui )] ≤ M for some lower semi-continuous ϕi : Ui → R+ which satisfies limn→∞ inf ui ∈K i i
/ n ϕ (u ) = ∞. Then, an
optimal team policy exists.

Remark 10.14. Building on [351, Theorems 2.3 and 2.5] and [151, p. 1691] (due to Blackwell’s theorem on irrelevant
information [48, 51], [349, p. 457]), an optimal policy, when exists, can be assumed to be deterministic.
10.8 Existence of Optimal Solutions 233

So far, we presented existence results for static or dynamic teams that admit independent static reduction. In the following,
we present existence results for teams that do not admit independent static reduction.

Theorem 10.15. [351, Theorem 2.9] Consider a sequential team with a classical information QN structure with the further
property that σ(ω0 ) ⊂ σ(y 1 ) (under every policy, y 1 contains ω0 ). Suppose further that k=1 Uk is compact. If c is lower
semi-continuous and each of the kernels pn (defined in (10.47)) is weakly continuous so that
Z
f (y n ) pn (dy n |ω0 , y 1 , . . . , y n−1 , u1 , · · · , un−1 ) (10.50)

is continuous in ω0 , y 1 , · · · , y n−1 , u1 , · · · , un−1 for every continuous and bounded f , then there exists an optimal team
policy which is deterministic.

A further existence result along similar lines, for a class of static teams, is presented next.

Theorem 10.16. [346, Theorem 5.6] Consider a static team with a classical information Q structure (that is, with an ex-
N
panding information structure so that σ(y n ) ⊂ σ(y n+1 ), n ≥ 1). Suppose further that k=1 (Yk × Uk ) is compact.
If
c̃(y 1 , · · · , y N , u1 , · · · , uN ) := E[c(ω0 , u)|y, u]
is lower semi-continuous in u for every y, then there exists an optimal team policy which is deterministic.

Remark 10.17. The power of this last result may first seem limited. However, some reflection leads to the conclusion that,
in the continuous-time theory of stochastic control, a related but not identical argument has remarkable consequences. If
one makes the measurements independent via a change of measure argument, as in Girsanov’s celebrated argument, so
that the information structure is first made static, and then makes the information structure classical by considering the
actions at time t measurable on the filtration generated by the past noise processes and actions up to time t; the proof of
Theorem 10.16 can be slightly adapted to show that such a set of measurement-action measures (with fixed marginal on the
measurements) that satisfy conditional independence u[0,t] ↔ y[0,t] ↔ ys − yt is weakly closed (these are known as wide-
sense admissible control policies). Furthermore, the value is continuous in this joint measure on {(u, y)s , s ∈ [0, T ]} and
this set of measures is tight. These lead to the compactness-continuity conditions and accordingly an existence result for
optimal policies follows. Furthermore, by showing that the set of {(u, y)s , s ≥ 0} measures which have quantized support
in the measurement variable are dense, one can show also that piece-wise constant control policies are nearly optimal. This
allows one to approximate a continuous-time process with a (sampled) discrete-time process and the machinery developed
earlier in the lecture notes are applicable. This approach is the essence of Kushner’s method [202], though stated somewhat
differently.

10.8.1 Some Applications and Revisiting Existence Results for Classical (Single-DM) Stochastic Control

Witsenhausen’s counterexample with Gaussian variables

Consider the celebrated Witsenhausen’s counterexample [325] as depicted in Figures 10.3 and 10.2: This is a dynamic
non-classical team problem with y 1 and w1 zero-mean independent Gaussian random variables with unit variance and
u1 = γ 1 (y 1 ), u2 = γ 2 (u1 + w1 ) and the cost function c(ω, u1 , u2 ) = k(y 1 − u1 )2 + (u1 − u2 )2 for some k > 0.
Witsenhausen’s counterexample can be expressed, through a change of measure argument (also due to Witsenhausen) as
in (10.39).
Since the optimal policy for γ 2 (y 2 ) = E[u1 |y 1 ] and E[(E[u1 |y 1 ])2 ] ≤ E[(u1 )2 ], it is evident with a two-stage analysis (see
[151, p. 1701]) that without any loss we can restrict the policies to be so that E[(ui )2 ] ≤ M for some finite M , for i = 1, 2;

2 1
this ensures a weak compactness condition on both γ̂ 1 and γ̂ 2 . Since the reduced cost (k(u1 −y 1 )2 +(u1 −u2 )2 ) η(yη(y−u
2)
)

is continuous in the actions, Theorem 10.13 applies.


234 10 Decentralized and Multi-Agent Stochastic Control

Existence for partially observable Markov Decision Processes (POMDPs)

Consider a partially observable stochastic control problem (POMDP) with the following dynamics.

xt+1 = f (xt , ut , wt ), yt = g(xt , vt ).

Here, xt is the X-valued state, ut is the U-valued the control, yt is the Y-valued measurement process. In this section, we
will assume that these spaces are finite dimensional real vector spaces. Furthermore, (wt , vt ) are i.i.d noise processes and
{wt } is independent of {vt }. The controller only has causal access to {yt }: A deterministic admissible control policy Π is
a sequence of functions {γt } so that ut = γ(y[0,t] ; u[0,t−1] ). The goal is to minimize

T
X −1
ExΠ0 [ c(xt , ut )],
t=0

for some continuous and bounded c : X × U → R+ .


Such a problem can be viewed as a decentralized stochastic control problem with increasing information, that is, one with
a classical information structure.
Any POMDP can be reduced to a (completely observable) MDP [352], [262], whose states are the posterior state distribu-
tions or beliefs of the observer. A standard approach for solving such problems then is to reduce the partially observable
model to a fully observable model (also called the belief-MDP) by defining

πt (A) := P (xt ∈ A|y[0,t] , u[0,t−1] ), A ∈ B(X)

and observing that (πt , ut ) is a controlled Markov chain where πt is P(X)-valued with P(X) being the space of probability
measures on X under the weak convergence topology. Through such a reduction, existence results can be established by
obtainingR conditions which would ensure that the controlled Markovian kernel for the belief-MDP is weakly continuous,
that is if F (πt+1 )P (dπt+1 |πt = π, ut = u) is jointly continuous (weakly) in π and u for every continuous and bounded
function F on P(X).
This was studied recently in [127, Theorem 3.7, Example 4.1] and [183] (see also [72] in a control-free context). In the
context of the example presented, if f (·, ·, w) is continuous and g has the form: yt = g(xt ) + vt , with g continuous and
wt admitting a continuous density function η, an existence result can be established building on the measurable selection
criteria under weak continuity [164, Theorem 3.3.5, Proposition D.5], provided that U is compact.
On the other hand, through Theorem 10.16, such an existence result can also be established by obtaining a static reduction
under the aforementioned conditions. Indeed, through (10.31), with η denoting the density of vn , we have P (yn ∈ B|xn ) =
− g(xn ))dy. With η and g continuous and bounded, taking y n := yn , by writing xn+1 = f (xn , un , wn ) =
R
B
η(y
f (f (xn−1 , un−1 , wn−1 ), un , wn ), and iterating inductively to obtain

xn+1 = hn (x0 , u[0,n−1] , w[0,n−1] ),

for some hn which is continuous in u[0,n−1] for every fixed x0 , w[0,n−1] , one obtains a reduced cost (10.31) that is a contin-
uous function in the control actions. Theorem 10.16 then implies the existence of an optimal control policy.R This reasoning
is also applicable when the measurements are not additive in the noise but with P (yn ∈ B|xn = x) = B m(y, x)η(dy)
for some m continuous in x and η a reference measure.

Revisiting fully observable Markov Decision Processes with the construction presented in the chapter

Consider a fully observed Markov decision process where the goal is to minimize
T
X −1
ExΠ0 [ c(xt , ut )],
t=0
10.9 Approximation of Optimal Solutions via Finite Approximations 235

for some continuous and bounded c : X × U → R+ . Suppose that the controller has access to x[0,t] , u[0,t−1] at time t. This
system can always be viewed as a sequential team problem with a classical information structure. Under the assumption
that the transition kernel according to the usual formulation, that is P (dx1 |x0 = x, u0 = u) is weakly continuous (in
the sense discussed in the previous application above), it follows that the transition kernel according to the formulation
introduced in (10.47) is also weakly continuous by an application [287, Theorem 3.5]. It follows that when U is compact,
and hence the existence of an optimal policy follows. A similar analysis is applicable when one considers the case where
P (dx1 |x0 = x, u0 = u) is strongly continuous in u for every fixed state x and the bounded cost function is continuous
only in u (this is another typical setup where measurable selection conditions hold (see Assumptions 5.2.1 and 5.2.2)).

10.9 Approximation of Optimal Solutions via Finite Approximations

In this section, we consider the finite approximation of static team problems. Since results of this section can also be applied
to static reduction of dynamic teams, we suppose that the cost function c also depends on the measurements y (which is
not the case in the original problem formulation). Recall that, in the independent static reduction of a dynamic team, the
reduced cost function cs is a function of ω0 , u, and y. To obtain finite approximation result, the following assumptions are
imposed on the components of the model.

Assumption 10.9.1 (a) The cost function c is continuous in (u, y) for any fixed ω0 . In addition, it is bounded on any
QN
compact subset of Ω0 × k=1 (Yk × Uk ).
(b) For each k, Uk is a closed and convex subset of a completely metrizable locally convex vector space.
(c) For each k, Yk is locally compact.
QN
(d) For any subset G of k=1 Uk , the function wG (ω0 , y) := supu∈G c(ω0 , y, u) is integrable with respect to
QN QN
µ(dω0 , dy), for any compact subset G of k=1 Uk of the form G = k=1 Gk .
(e) For any γ ∈ Γ with J(γ) < ∞ and each k, there exists uk,∗ ∈ Uk such that J(γ −k , γukk,∗ ) < ∞, where γukk,∗ ≡
uk,∗ .
QN
In what follows, for any subset G of k=1 Uk , we let
 N
! 
Y
k
ΓG := γ∈Γ:γ Y ⊂G
k=1

and Γc,G := Γc ∩ ΓG , where Γc denotes the set of continuous strategies. Using these definitions, let us define the following
QN
set of strategic measures for any subset G of k=1 Uk :

LG
A (µ)
 Z N
Y 
:= P ∈ LA (µ) : P (B) = QN µ(dω0 , dy) 1{uk =γ k (yk )∈B k } , γ ∈ ΓG .
B0 × Ak k=1
k=1

Let LG,c G
A (µ) denote the set of strategic measures in LA (µ) induced by continuous policies.

The following
QN result states that, there exists a near optimal strategic measure whose support on the product of action
spaces k=1 Uk is convex and compact (and thus bounded) subset G of it, and conditional distributions of actions given
measurements are induced by continuous policies.
QN
Proposition 10.18. Suppose Assumption 10.9.1 holds. Then, for any ε > 0 there exists a compact subset G of k=1 Uk of
QN
the form G = i=1 Gi , where each Gi is convex and compact, such that
Z
inf P (ds) c(s) < J ∗ + ε.
P ∈LG,c
A
(µ)
236 10 Decentralized and Multi-Agent Stochastic Control

Given any strategic measure, using Assumption 10.9.1-(e) and the fact that every measure on a Borel space is tight [249,
Theorem 3.2], one can construct a strategic measure in LA (µ) whose support on the product of action spaces is convex
and compact and whose cost is ε/2-close to the cost of the given strategic measure. For the new strategic measure, since
it has a convex and compact support on the product of action spaces, using Lusin’s theorem [110, Theorem 7.5.2], we
can construct a strategic measure induced by continuous policies whose cost function is ε/2-close to the cost of bounded
support strategic measure. We can complete the proof by combining these two results.

Since each Yi is a locally compact S separable metric space, there exists an increasing sequence of compact subsets {Kli }

such that Kli ⊂ int Kl+1
i
and Yi = l=1 Kli [4, Lemma 2.76], where int D denotes the interior of the set D.

Let di denote the metric on Yi . For each l ≥ 1, let Yil,n := yi,1 , . . . , yi,il,n ⊂ Kli be an 1/n-net in Kli . Recall that if
Yil,n is an 1/n-net in Kli , then for any y ∈ Kli we have

1
min di (y, z) < .
z∈Yl,ni n

i
For each l and n, let ql,n : Kli → Yil,n be a nearest neighborhood quantizer given by
i
ql,n (y) = arg min
i
di (y, z),
z∈Yl,n

i
where ties are broken so that ql,n is measurable. If Kli = [−M, M ] ⊂ Yi = R for some M ∈ R+ , the finite set Yil,n can
be chosen such that ql,n becomes a uniform quantizer. We let Qil,n : Yi → Yil,n denote the extension of ql,n
i i
to Yi given by
(
i
ql,n (y), if y ∈ Kli ,
Qil,n (y) :=
yi,0 , otherwise,

where yi,0 ∈ Yi is some auxiliary element.


i
Define Γl,n = Γ i ◦ Qil,n ⊂ Γ i ; that is, Γl,n
i
is defined to be the set of all strategies γ̃ i ∈ Γ i of the form γ̃ i = γ i ◦ Qil,n ,
N
where γ i ∈ Γ i . Define also Γl,n := i=1 Γl,n i i
Q
⊂ Γ. Note that, for any i = 1, . . . , N , Γl,n is the set of policies for DM i
which can only use the output levels of the quantizer Ql,n . In other words, in addition to measurement channel g i (dy i |ω0 )
i

between DM i and the Nature, there is also an analog-to-digital converter (quantizer) between them.
Using these definitions, let us define the following set of strategic measures for any l and n:

Ll,n
A (µ)
 Z N
Y 
:= P ∈ LA (µ) : P (B) = QN µ(dω0 , dy) 1{uk =γ k (yk )∈B k } , γ ∈ Γl,n .
B0 × Ak k=1
k=1

The following theorem states that an optimal (or almost optimal) strategic measure can be approximated with arbitrarily
small approximation error for the induced costs by strategic measures in Ll,n
A (µ) for sufficiently large l and n.

QN QN
Theorem 10.19. [274] For any ε > 0, there exist (l, n(l)), a compact subset G of k=1 Uk of the form G = i=1 Gi ,
l,n(l)
where each Gi is convex and compact, and P ∈ LA (µ) LG
T
A (µ) such that
Z
P (ds) c(s) < J ∗ + ε

For each (l, n), we define a team model with finite measurement spaces. We prove that, for sufficiently large l and n,
optimal strategic measure of the team model corresponding to (l, n) will provide a strategic measure to the original team
model which is nearly optimal.
10.9 Approximation of Optimal Solutions via Finite Approximations 237

To this end, fix any (l, n). For the pair (l, n), the corresponding finite measurement team model has the following measure-
ment spaces: Zil,n := {yi,0 , yi,1 , . . . , yi,il,n } (i.e., the output levels of Qil,n ), i ∈ N . The stochastic kernels gl,n
i
( · |ω0 ) from
i
Ω0 to Zl,n denotes the measurement constraints and given by:

il,n
X l,n
i
gl,n ( · |ω0 ) := g(Si,j |ω0 ) δyi,j ( · ),
j=0

l,n 
where Si,j := y ∈ Yi : Qil,n (y) = yi,j . Indeed, gl,n
i
( · |ω0 ) is the push-forward of the measure g i ( · |ω0 ) with respect to
i
the quantizer Ql,n .
 QN
Let Φin,l := ϕi : Zil,n → Ui , ϕi measurable denote the set of measurable policies for DM i and let Φl,n := i=1 Φil,n .
The cost of this team model is Jl,n : Φl,n → R+ and defined as
Z
Jl,n (ϕ) := QN c(ω0 , y, u) Pl,n (dω0 , dy),
Ω0 × Zil,n
i=1

where ϕ = (ϕ1 , . . . , ϕN ), u = ϕ(y), and

N
Y
i
Pl,n (dω0 , dy) = P (dω0 ) gl,n (dy i |ω0 ) =: µl,n (dω0 , dy).
i=1

QN QN
For any compact subset G of k=1 Uk , we also define ΦG i
l,n := {ϕ ∈ Φl,n : ϕ( i=1 Zl,n ) ⊂ G}.

In order to obtain the approximation result, we need to impose the following additional assumption.
QN QN
Assumption 10.9.2 For any compact subset G of k=1 Uk of the form G = i=1 Gi , we assume that the function wG is
uniformly integrable with respect to the measures {µl,n }; that is,
Z
lim sup wG (ω0 , y) dµl,n = 0.
R→∞ l,n {wG >R}

Note that Assumption 10.9.1-(d),(e) hold if the cost function is bounded. Indeed, conditions in Assumption 10.9.1 are quite
mild and hold for the celebrated counterexample of Witsenhausen.

Theorem 10.20. [274] Suppose Assumptions 10.9.1 and 10.9.2 hold. Then, for any ε > 0, there exists a pair (l, n(l)) and
QN QN
a compact subset G = i=1 Gi of k=1 Uk such that an optimal (or almost optimal) strategic measure P l,n(l) in the set
QN
TAG (µl,n(l) ) for the (l, n(l)) team is ε-optimal for the original team problem when P l,n(l) is extended to Ω× k=1 (Yk ×Uk )
via quantizers Qil,n(l) ; that is,

Z N
Y
l,n(l)
Pex (·) = µ(dω0 , dy) 1{uk =γ k ◦Qk (y k )∈·}
l,n(l)
· k=1

where
Z N
Y
P l,n(l) (B) = µl,n(l) (dω0 , dy) 1{uk =γ k (yk )∈·}
· k=1
238 10 Decentralized and Multi-Agent Stochastic Control

10.10 Dynamic Programming and Centralized MDP Reduction Approaches to Team Decision
Problems

10.10.1 Dynamic programming approach based on Common Information and a Controlled Markov State

In a team problem, if all the random information at any given decision maker is common knowledge between all decision
makers, then the system is essentially centralized. If only some of the system variables are common knowledge, the re-
maining unknowns may or may not lead to a computationally tractable program generating an optimal solution. A possible
approach toward establishing a tractable program is through the construction of a controlled Markov chain where the con-
trolled Markov state may now live in a larger state space (for example a space of probability measures) and the actions are
elements in possibly function spaces. This controlled Markov construction may lead to a computation of optimal policies.
Such a dynamic programming approach has been adopted extensively in the literature (see for example, [19], [338], [86],
[3], [342], and generalized in [241,242]) through the use of a team-policy which uses common information to generate par-
tial functions for each DM to generate their actions using local information. Thus, in the dynamic programming approach,
a separation of team decision policies in the form of a two-tier architecture, a higher-level controller and a lower-level
controller, can be established with the use of common knowledge.
In the following, we present the ingredients of such an approach, as generalized in [242] and termed the common informa-
tion approach:
1. Elimination of irrelevant information at the DMs: In this step, irrelevant local information at the DMs, say DM k,
is identified as follows. By letting the policy at other DMs to be arbitrary, the policy of DM k can be optimized as
a best-response function, and irrelevant data at DM k can be removed.
2. Construction of a coordinated system: This step identifies the common information and local/private information
at the DMs, after Step 1 above has been carried out. A fictitious coordinator (higher-level controller) uses the
common information to generate team policies, which in turn dictates the (lower-level) DMs what to do with their
local information.
3. Formulation of the cost function as a Partially Observed Markov Decision Process (POMDP), in view of the coor-
dinator’s optimal control problem: A fundamental result in stochastic control is that the problem of optimal control
of a partially observed Markov chain (with additive per-stage costs) can be solved by turning the problem into a
fully observed one on a larger state space where the state is replaced by the “belief” on the state.
4. Solution of the POMDP leads to the structural results for the coordinator to generate optimal team policies, which
in turn dictates the DMs what actions to take given their local information realizations.
5. Establishment of the equivalence between the solution obtained and the original problem, and translation of the
optimal policies. Any coordination strategy can be realized in the original system. Note that, even though there is
no real coordinator, such a coordination can be realized implicitly, due to the presence of common information.
We will provide a further explicit setting with such a recipe at work, in the context of the k-stage periodic belief sharing
pattern in the next section. In particular, Lemma 10.10.1 and Lemma 10.10.2 will highlight this approach. When a given
information structure does not allow for the construction of a controlled Markov chain even in a larger, but fixed for all
time stages, state space, one question that can be raised is what information requirements would lead to such a structure.
We will also investigate this problem in the context of the one-stage belief sharing pattern in the next section.

k-Stage Periodic Information or Belief Sharing Pattern

In this section, we will use the term belief for a probability measure-valued random variable. This terminology has been
used particularly in the artificial intelligence and computer science communities, which we adopt here. We will, however,
make precise what we mean by such a belief process in the following.
As mentioned earlier in Chapter 6, a fundamental result in stochastic control is that the problem of optimal control of a
partially observed Markov chain can be solved by turning the problem into a fully observed one on a larger state space
10.10 Dynamic Programming and Centralized MDP Reduction Approaches to Team Decision Problems 239

where the state is replaced by the belief on the state. Such an approach is very effective in the centralized setting; in
a decentralized setting, however, the notion of a state requires further specification. In the following, we illustrate this
approach under the k-step periodic belief sharing information pattern.
Consider a joint process {xt , yt , t ∈ Z+ }, where we assume for simplicity that the spaces where xt , yt take values from
are finite dimensional real-valued or countable. They are generated by

xt+1 = f (xt , u1t , . . . , uL


t , wt ),

yti = g(xt , vti ),


where xt is the state, uit ∈ Ui is the control action, (wt , vti , 1 ≤ i ≤ L) are second order, zero-mean, mutually independent,
i.i.d. noise processes. We also assume that the state noise, wt , either has a probability mass function, or a probability
measure with a density function.
Suppose that there is a common information vector Itc at some time t, which is available to all the decision makers. At
c
times ks−1, with k > 0 fixed, and s ∈ Z+ , the decision makers share all their information: Iks−1 = {y[0,ks−1] , u[0,ks−1] }
c
and for I0 = {P (x0 )}, that is at time 0 the DMs have the same a priori belief on the initial state. Hence, at time t, DM i
i c
has access to {y[ks,t] , Iks−1 }.

Until the next common observation instant t = k(s + 1) − 1 we can regard the individual decision functions specific to DM
i
i as {uit = γsi (y[ks,t] c
, Iks−1 )}; we let γs denote the ensemble of such decision functions and let γ denote the team policy.
i
It then suffices to generate γs for all s ≥ 0, as the decision outputs conditioned on y[ks,t] i
, under γsi (y[ks,t] c
, Iks−1 ), can
c c
be generated. In such a case, we can define γs (., Iks−1 ) to be the joint team decision rule mapping Iks−1 into a space of
i
action vectors: {γsi (y[ks,t] c
, Iks−1 ), i ∈ L = {1, 2 . . . , L}, t ∈ {ks, ks + 1, . . . , k(s + 1) − 1}}.

Let [0, T − 1] be the decision horizon, where T is divisible by k. Let the objective of the decision makers be the joint
minimization of
T −1
1 2 L X
Exγ0 ,γ ,...,γ [ c(xt , u1t , u2t , . . . , uL
t )],
t=0
1 2 L
over all policies γ , γ , . . . , γ , with the initial condition x0 specified. The cost function
T −1
γ X
Jx0 (γ) = Ex0 c(xt , ut )
t=0

can be expressed as:


T
k −1
γ X c
Jx0 (γ) = Ex0 [ c̄(γs (., Iks−1 ), x̄s )]
s=0

with
k(s+1)−1
c γ X
c̄(γs (., Iks−1 ), x̄s ) = Ex̄s [ c(xt , ut )]
t=ks

Lemma 10.10.1 [342] Consider the decentralized system setup above. Let Itc be a common information vector supplied to
the DMs regularly every k time stages, so that the DMs have common memory with a control policy generated as described
c
above. Then, {x̄s := xks , γs (·, Iks−1 ), s ≥ 0} forms a controlled Markov chain.

In view of the above, we have the following separation result.

Lemma 10.10.2 [342] Let Itc be a common information vector supplied to the DMs regularly every k time steps. There is
c c
no loss in performance if Iks−1 is replaced by P (x̄s |Iks−1 ).
240 10 Decentralized and Multi-Agent Stochastic Control

An essential issue for a tractable solution is to ensure a common information vector which will act as a sufficient statistic
for future control policies. This can be done via sharing information at every stage, or some structure possibly requiring
larger but finite delay.
The above motivates us to introduce the following pattern.

Definition 10.10.1 k-stage periodic belief sharing pattern [342] An information pattern in which the decision makers
share their posterior beliefs to reach a joint belief about the system state is called a belief sharing information pattern. If
the belief sharing occurs periodically every k-stages (k > 1), the DMs also share the control actions they applied in the
last k − 1 stages, together with intermediate belief information. In this case, the information pattern is called the k-stage
periodic belief sharing information pattern.

Remark 10.21. For k > 1, it should be noted that, the exchange of the control actions is essential. ⋄

The above generalize to models with standard Borel spaces [239], where weak Feller regularity are also obtained for the
reduced model. Accordingly, the numerical and learning theoretic results are applicable.

10.10.2 A Universal Dynamic Program

[346] considered the following topology on control policies, while developing a universal dynamic programming algorithm
applicable to any sequential decentralized stochastic control problem, generalizing Witsenhausen’s program [327] which
is tailored primarily for countable probability spaces.
Define
(i) State: xt = {ω0 , u1 , · · · , ut−1 , y 1 , · · · , y t }, 1 ≤ t ≤ N .
Qt Qt−1 Qt Qt−1
(i’) Extended State: πt ∈ P(Ω0 × i=1 Yi × i=1 Ui ) where, for Borel B ∈ Ω0 × i=1 Yi × i=1 Ui ,

πt (B) := Eπt [1{(ω0 ,y1 ,··· ,yt ;u1 ,··· ,ut−1 )∈B} ].
Qt Qt−1
Thus, πt ∈ P(Ω0 × i=1 Yi × i=1 Ui ) where the space of probability measures is endowed with the weak
convergence topology.
Qt Qt
(ii) Control Action: Given πt , γ̂ t is a probability measure in P(Ω0 × k=1 Yk × k=1 Uk ) that satisfies the conditional
independence relation:
ut ↔ y t ↔ xt = (ω0 , y 1 , · · · , y t ; u1 , · · · , ut−1 )
(that is, for every Borel B ∈ Ui , almost surely under γ̂ t , the following holds:

P (ut ∈ B|y t , (ω0 , y 1 , · · · , y t ; u1 , · · · , ut−1 )) = P (ut ∈ B|y t )

with the restriction


xt ∼ πt .
Denote with Γ t (πt ) the set of all such probability measures. Any γ̂ t ∈ Γ t (πt ) defines, for almost every realization
y t , a conditional probability measure on Ut . When the notation does not lead to confusion, we will denote the action
at time t by γ t (dut |y t ), which is understood to be consistent with γ̂ t .
(ii’)Alternative Control Action for Static Teams with Independent Measurements: Given πt , γ̂ t is a probability measure
t t
on Yt × Ut with a fixed marginal P (dy t ) on Yt , that is πtY (dy t ) = P (dy t ). Denote with Γ t (πtY ) the set of all
such probability measures. As above, when the notation does not lead to confusion, we will denote the action at
time t by γ t (dut |y t ), which is understood to be consistent with γ̂ t . In particular, (y t , ut ) is independent of (y k , uk )
for k ̸= t.
With the control actions defined as in the above [346] developed a universal dynamic program for any sequential decen-
tralized stochastic control and established, as a corollary of the program, further existence results, one which is essentially
10.12 Exercises 241

identical to that presented in 10.13, but slightly more restrictive in that the cost function is assumed to be continuous in all
of its arguments.

Theorem 10.22. [346]


(i) Under the kernel (10.47) and controlled Markov construction presented, the optimal team problem admits a well-
defined backwards-induction (dynamic programming) recursion.
(ii) In particular, if the problem is independent static-reducible, actions are compact-valued and the cost function is
continuous, an optimal policy exists and the value function is continuous in the prior (that is, in the distribution of
primitive noise variables) under weak convergence.

Remark 10.23. The above construction is related to an interpretation put forward by Witsenhausen in his standard form
[327] where all the uncertainly is embedded into the initial state and the controlled system evolves deterministically.
Witsenhausen had considered only countable probability spaces for an optimality analysis.

Remark 10.24. The fully observed MDP setup can be viewed as a special case of the above. In this context, by Blackwell’s
theorem (Theorem 5.1.1), we know that we can reduce the search space to policies that are Markov. In this case, the
optimality analysis via Bellman’s principle (Theorem 5.1.3) can be recovered via the Universal Dynamic Program.

10.11 Bibliographic Notes

We primarily followed [346], [351] and Chapters 2, 3, 4 and 12 of [349] for this topic. For a more complete coverage, the
reader may follow [349].
In the economics and game theory literature, information structures are also studied extensively. Stochastic team problems
are termed as identical interest games. In this literature, LC (µ) appears in the analysis of Aumann’s correlated equilibrium
[23]. Common and independent randomness discussions appear in the analysis of comparison of information structures
[215]. For further discussions, including a multi-stage generalization known as communication equilibria, see [134]. For a
detailed treatment, we refer the reader to [229, p. 131].

10.12 Exercises

Exercise 10.12.1 Consider the following static team decision problem with dynamics:

x1 = ax0 + b1 u10 + b2 u20 + w0 ,

y01 = x0 + v01 ,
y02 = x0 + v02 ,
Here v01 , v02 , w0 are independent, Gaussian, zero-mean with unit variance.
Let γ i : R → R be policies of the controllers: u10 = γ01 (y01 ), u20 = γ02 (y02 ).
Find 1 2
min
1 2
Eνγ0 ,γ [x21 + ρ1 (u10 )2 + ρ2 (u20 )2 ],
γ ,γ

where ν0 is a zero-mean Gaussian distribution and ρ1 , ρ2 > 0.


a) Find an optimal team policy γ = {γ 1 , γ 2 }.
b) When b1 = b2 and ρ1 = ρ2 , can you conclude that an optimal solution will be identical for both Decision Makers?
See [278–280] for further structural results on convex and exchangeable teams.
242 10 Decentralized and Multi-Agent Stochastic Control

Exercise 10.12.2 Consider the following team decision problem with dynamics:

xt+1 = axt + b1 u1t + b2 u2t + wt ,

yt1 = xt + vt1 ,
yt2 = xt + vt2 ,

Here x0 , vt1 , vt2 , wt are mutually and temporally independent zero-mean Gaussian random variables.
Let {γ i } be the policies of the controllers so that uit = γti (y0i , y1i , · · · , yti ) for i = 1, 2.
Consider:
−1
 TX  
1 2
min
1 2
Exγ0 ,γ x2t + ρ1 (u1t )2 + ρ2 (u2t )2 ] + x2T ,
γ ,γ
t=0

where ρ1 , ρ2 > 0.
Explain if the following are correct or not:
a) For T = 1, the problem is a static team problem.
b) For T = 1, optimal policies are linear.
c) For T = 1, linear policies may be person-by-person-optimal. That is, if γ 1 is assumed to be linear, then γ 2 is linear;
and if γ 2 is assumed to be linear then γ 1 is linear.
d) For T = 2, optimal policies are linear.
e) For T = 2, linear policies may be person-by-person-optimal.

Exercise 10.12.3 Consider a common probability space (with a finite sample space Ω) on which the information available
to two decision makers DM1 and DM2 are defined, such that I1 is available at DM1 and I2 is available at DM2 .
R. J. Aumann [22] defines that an information E is common knowledge between two decision makers DM1 and DM2 , if
whenever E happens, DM1 knows E, DM2 knows E, DM1 knows that DM2 knows E, DM2 knows that DM1 knows E, and
so on.
Let Ω be finite. Suppose that one claims that an event E is common knowledge if and only if E ∈ σ(I1 ) ∩ σ(I2 ), where
σ(I1 ) denotes the σ−field over Ω generated by information I1 and likewise for σ(I2 ).
Is this argument correct? Provide an answer with precise arguments. You may wish to consult [22], [244], [70] and Chapter
12 of [349].

Exercise 10.12.4 Let X be a binary random variable. Suppose two decision makers DM 1 and DM 2 have access to some
local random variables Y 1 and Y 2 , respectively, defined on a common probability space and correlated with X, and
exchange their conditional expectations over time. Suppose further that:
– the information σ-fields at each decision maker is increasing: Fti ⊂ Ft+1
i
, i = 1, 2, t ∈ Z+ .
i
– for all n ∈ N, there exists m > n such that Fm contains information on E[X|Fnj ], i, j = 1, 2. That is, the decision
makers exchange their estimates (but not their raw data -Y i is private to DM i, i = 1, 2-) infinitely often.
State and rigorously justify your answers for the following:
a) [10 Points] Is there a limit for limn→∞ E[X|Fnj ], j = 1, 2? Either argue that the limit exists, or provide a counterex-
ample.
b) [10 Points] For the cases where the limit exists, is it the case that

lim E[X|Fn1 ] = lim E[X|Fn2 ]


n→∞ n→∞
10.12 Exercises 243

Either prove the result, or provide a counterexample.


Hint: See [65] (see also [139] and [307])

Exercise 10.12.5 Consider a linear Gaussian system with mutually independent and i.i.d. noises:
L
X
xt+1 = Axt + B j ujt + wt ,
j=1

yti = C xt + vti , 1 ≤ i ≤ L,
i
(10.51)

with the one-step delayed observation sharing pattern.


Construct a controlled Markov chain for the team decision problem: First show that one could have

{yt1 , yt2 , . . . , ytL , P (dxt |y[0,t−1]


1 2
, y[0,t−1] L
, . . . , y[0,t−1] )}

as the state of the controlled Markov chain.


Consider the following problem:
T −1
γ X
Eν0 [ c(xt , u1t , · · · , uL
t )]
t=0
1 2 L
For this problem, if at time t ≥ 0 each of the decision makers (say DM i) has access to P (dxt |y[0,t−1] , y[0,t−1] , . . . , y[0,t−1] )
i
and their local observation y[0,t] , show that they can obtain a solution where the optimal decision rules only uses
1 2 L
{P (dxt |y[0,t−1] , y[0,t−1] , . . . , y[0,t−1] ), yti }:
1 2 L i
What if, they do not have access to P (dxt |y[0,t−1] , y[0,t−1] , . . . , y[0,t−1] ), and only have access to y[0,t] ? What would be a
sufficient statistic for each decision maker for each time stage?

Exercise 10.12.6 Two decision makers, Alice and Bob, wish to control a system:

xt+1 = axt + uat + ubt + wt ,

yta = xt + vta ,
ytb = xt + vtb ,
where uat , yta are the control actions and the observations of Alice, ubt , ytb are those for Bob and vta , vtb , wt are independent
zero-mean Gaussian random variables with finite variance. Suppose the goal is to minimize for some T ∈ Z+ :
−1
 TX 
a
,Π b
ExΠ0 x2t + ra (uat )2 + rb (ubt )2 ,
t=0

for ra , rb > 0, where Π a , Π b denote the policies adopted by Alice and Bob. Let the local information available to Alice be
Ita = {ysa , uas , s ≤ t − 1} ∪ {yta }, and Itb = {ysb , ubs , s ≤ t − 1} ∪ {ytb } is the information available at Bob at time t.
Consider an n−step delayed information pattern: In an n−step delayed information sharing pattern, the information at
Alice at time t is
Ita ∪ It−n
b
,
and the information available at Bob is
Itb ∪ It−n
a
.
State if the following are true or false:
a) If Alice and Bob share all the information they have (with n = 0), it must be that, the optimal controls are linear.
244 10 Decentralized and Multi-Agent Stochastic Control

b) Typically, for such problems, for example, Bob can try to send information to Alice to improve her estimation on the
state, through his actions. When is it the case that Alice cannot benefit from the information from Bob, that is for what
values of n, there is no need for Bob to signal information this way?
c) If Alice and Bob share all information they have with a delay of 2, then their optimal control policies can be written as

uat = fa (E[xt |It−2


a b
, It−2 a
], yt−1 , yta ),

ubt = fb (E[xt |It−2


a b
, It−2 b
], yt−1 , ytb ),
for some functions fa , fb . Here, E[.|.] denotes the expectation.
d) If Alice and Bob share all information they have with a delay of 0, then their optimal control policies can be written as

uat = fa (E[xt |Ita , Itb ]),

ubt = fb (E[xt |Ita , Itb ]),


for some functions fa , fb . Here, E[.|.] denotes the expectation.
11

Controlled Stochastic Differential Equations

This chapter introduces the basics of stochastic differential equations and then studies controlled such equations. A com-
plete treatment is beyond the scope of these notes, however, the essential tools and ideas will be presented so that a student
who is comfortable with the discrete-time discussion thus far in the notes can realize that with a little additional effort
the continuous-time case can also be followed with ease. The reader is referred to e.g. [15, 174, 180, 203, 248] for more
comprehensive treatments on various aspects ranging from mathematical foundations, stability, optimal control, filtering,
and numerical methods.
Our approach here will primarily be to map the material presented so far in the notes to the continuous-time case, with the
understanding that the discrete-time theory is well understood. With Xt an R-valued random variable for each t ∈ R+ ,
consider a stochastic process Xt , t ∈ R+ . Given a sufficiently regular function f suppose that we can define

E[f (Xh )|X0 = x] − f (x)


lim =: Af (x), x∈R
h→0 h
o(h)
for some map A (to be studied further). This means that E[f (Xh )|X0 = x] = f (x) + Af (x)h + o(h), where h → 0 as
h → 0. With µt (B) = E[1{Xt ∈B} ], for all Borel B, the above implies under mild conditions on A that
Z Z tZ  Z
µt (dx)f (x) = Af (z)µs (dz) ds + f (x)µ0 (dx)
0

We will observe that the above can be viewed as a limit (as h → 0) of interpolations of the sampled (and thus discrete with
k ∈ Z+ ) stochastic process

X(k+1)h = Xkh + hb(Xkh ) + σ(Xkh ) hZ (11.1)
2
where Z ∼ N (0, 1) and Af (x) = b(x) dx d
f (x) + 12 (σ 2 (x)) ∂∂xf2 (x). In the limit as h → 0, we arrive in some particular
sense (that of weak convergence of path valued random processes under the topology of uniform convergence over compact
sets), at the limit equation
dXt = b(Xt ) + σ(Xt )dBt ,
which is called a stochastic differential equation. Here, Bt is the Brownian motion.
The discussion (11.1) also leads to the following chain rule: Let f (x, t) be differentiable so that the operations to follow
are well-defined (e.g., twice continuously differentiable in x and continuously differentiable in t): Then, if we attempt to
write
∂f ∂f ∂f ∂f √
f (xt+h , t + h) ≈ f (x, t) + h+ dx = f (x, t) + h+ (b(x)h + σ(x) hZ)
∂t ∂x ∂t ∂x
∂f ∂f

what we observe is that in the last term ∂x dx = ∂x (b(x)h + σ(x) hZ), when normalization by h is made, the expression

h/h does not decay to zero and and the second derivative term appearing in the Taylor’s expansion, which would be
2
1∂ f
2 ∂x2 (dx)(dx), is non-negligible. Accordingly, a more appropriate expression is:
246 11 Controlled Stochastic Differential Equations

∂f ∂f 1 ∂2f
f (xt+h , t + h) ≈ f (x, t) + h+ dx + (dx)(dx)
∂t ∂x 2 ∂x2
leading to
∂f ∂f √ 1 ∂2f 2
f (xt+h , t + h) ≈ f (x, t) + h+ (hb(xt ) + σ(xt ) hZ) + σ (xt )hZ 2
∂t ∂x 2 ∂x2
This essentially leads to Itô’s formula to be studied. A number of technical questions will arise with respect to the notion of
convergence as h ↓ 0 and the non-differentiability of the Brownian process Bt . This model will be generalized, and there
will also be control entering the flow, e.g. via b(x, u) with u denoting the control term and possibly in σ(·) as well.
We will restrict the model to certain systems, e.g. those driven by the Brownian process, though one can in principle study
more general models (the term multiplying σ(Xkh ) does not need to be a Gaussian measure and there exist many other
processes that can be considered.
We should note that the construction of a stochastic process on a continuous time interval, such as [0, T ] requires more
caution when compared with a discrete-time stochastic process, as we will observe. In this chapter, we will primarily be
concerned with controlled Markov processes Xt , each taking values in Rn for t ∈ [0, T ] or t ∈ [0, ∞) and where the
integration term involves the Brownian process or semimartingale processes [174].

11.1 Continuous-time Markov processes

11.1.1 Two ways to construct a continuous-time Markov process

As discussed in Chapter 1 and Section 1.4, one way to define a stochastic process is to view it as a vector valued random
variable. This requires us to place a proper topology on the set of sample paths, to be discussed further below.
Another definition would involve defining the process on finitely many time instances: Let {Xt (ω), t ∈ [0, T ]} be stochastic
process so that for each t, Xt (ω) is an Rn -valued random variable measurable on some probability space (Ω, F, P ). We
can define the σ-algebra generated by cylinder sets (as in Chapter 1) of the form:

{ω ∈ Ω : Xt1 (ω) ∈ A1 , Xt2 (ω) ∈ A2 , · · · , XtN (ω) ∈ AN , Ak ∈ B(Rn ), N ∈ N}

By defining a stochastic process in this fashion and assigning probabilities to such finite dimensional events, Theorem 1.2.3
implies that there exists a unique stochastic process on the σ-algebra generated by the sets of this form. However, unlike a
discrete-time stochastic process, in general, not all properties of the stochastic process are captured by finite dimensional
distributions of it and the σ-field generated by such sets is not a sufficiently rich set of sets. For example the set of sample
paths that satisfy supt∈[0,1] |Xt (ω)| ≤ 10 may not be a well-defined event (that is, a set) in this σ-algebra. Likewise, the
extension theorem considered in Theorem 1.2.2 requires a probability measure already defined on the cylinder sets; it may
not be possible to define such a probability measure by only considering finite dimensional distributions [332].
If a stochastic process has continuous sample paths, then by specifying the process on rational time instances will uniquely
define the process. Thus, if the process is known to admit certain regularity properties, the technical issues with regard to
defining a process on finitely many sample points will disappear.

11.1.2 The Brownian motion

Definition 11.1.1 A stochastic process Bt is called a Wiener process or Brownian motion if (i) the finite dimensional
distributions of Bt are such that for any n ∈ N and any sequence 0 < t1 < t2 < · · · , tn , the collection of random variables
Bt1 , Bt2 − Bt1 , · · · , Btn − Btn−1 are independent Gaussian zero mean random variables with Bk − Bs ∼ N (0, k − s),
and (ii) Bt has continuous sample paths.

Such a process exists and can be constructed as a limit of random walks as briefly suggested in Remark 11.1 below. Going
back to the construction we discussed in the previous section, we can define the Brownian motion as a C([0, ∞))-valued
11.1 Continuous-time Markov processes 247

(that is, a continuous path valued) random variable: The topology on C([0, ∞)) is the topology of uniform convergence
on compact sets (this is a stronger convergence than the topology of point-wise convergence but weaker than the topology
of uniform convergence over R). This is in agreement with the finite dimensional characterization through which we could
define the Brownian motion.

Remark 11.1. [Why Brownian Motion?] The Gaussian property of the continuous limit process is universal in the sense
that, any continuous time process with sufficiently regular independent increments must be the Brownian process (via a
result known as Donsker’s theorem). In particular, even though typically in the construction of the Brownian motion (or
its existence), one considers Gaussian i.i.d. random increments and takes its limit; this is not necessary for the Gaussian
properties of the limit: Let {Z1 , Z2 , · · · , } be an i.i.d. random sequence with mean 0 and variance 1. For each n ∈ N define
the random variable (with variance t for each n ∈ N):
1 X
Wn (t) = √ Zk , t ∈ [0, 1]. (11.2)
n
1≤k≤⌊nt⌋

This is a random function. By the central limit theorem, Wn (t) → N (0, t) (in distribution, that is weakly) for each t ∈ R+
and the same holds for any finite collection of time values. With this insight, one can also show that the path-valued random
variable converges weakly (where one needs to define an appropriate metric on the path-valued realization space, as the
elements of the sequence may not be continuous) to the standard Brownian motion. In this context, an appropriate topology
is the Skorokhod topology defined on the space of functions which are right continuous with left limits: Such a topology
defines a separable metric space [42].
For many interesting properties of the Brownian motion, the reader is referred to [252].

Remark 11.2 (Going beyond the Brownian motion). While the discussion above justifies the typical usage of the Brownian
motion for many stochastic integration models studied later in the chapter, one can consider more general processes (known
as semimartingales) for the analysis in the following sections to be applicable [174]. Some applications may force one to
even be more general and consider driving signals that are to be studied under the theory of rough-paths [135, 220], which
seeks to give a sample path sense meaning to stochastic integration, to be discussed further, as well as several robustness
properties to approximate models and continuity of solutions in the driving noise.

On White Noise

In many physical systems, one encounters models of the form

dx
= (f (xt ) + ut ) + nt ,
dt
where nt is some noise process. In engineering, one would like to model the noise process to be white, in the sense that
nt , ns are independent for t ̸= s, and nt is zero-mean. We call such a process white, because the Fourier transform (and
thus the frequency spectrum) of the correlation function defined as R(τ ) = E[nt nt+τ ] of such a process is a constant:
If the process is a discrete-time process with finite support, then this interpretation would be directly applicable since the
Fourier transform of a discrete-time impulse would be constant for all frequency values. For a continuous-time process,
however, if R(τ ) = E[nt nt+τ ] = 0 for all τ ̸= 0, then a mathematical complication arises: If R(0) < ∞, then this signal
has zero-energy and its Fourier transform would be identically 0. If R(0) = ∞, then such an R would have significant
irregularities; such a process would have its correlation function as E[ns nt ] = δ(t − s); where the Dirac delta function
δ is a distribution acting on a proper set of test functions (such as the Schwartz signals S [173]). Such a process is not a
well-defined Gaussian process since it does not have a well-defined correlation function as δ itself is not a function. But,
one can view this process as a distribution, or always cautiously always work under an integral; this way one can make an
operational use for such a definition.
With such a cautious interpretation, as we see in Exercise 11.10.4, the Fourier transform of a Brownian motion, over a
bounded support, is i.i.d. across its discrete spectrum coefficients, and Gaussian. This justifies the term white noise. Thus,
it is evident that nt would not be a well-defined process and instead of nt , we will only work with its integral Bt . Thus,
248 11 Controlled Stochastic Differential Equations

while working with Bt , instead of derivatives, we will study integral equations. On the other hand, it will be evident that
we cannot take the ordinary Lebesgue or Riemann integrations for Bt (unless the function we work with is too regular)
since Bt is too irregular. Instead, a method to obtain integrations will be introduced: the Itô integral. The properties of
integration, differentiation, chain rule etc. for such integrations is called stochastic calculus. Later in the chapter, we will
add control to the dynamics.

11.2 Stochastic Integration, the Itô Integral and Stochastic Differential Equations

11.2.1 Some subtleties on stochastic integration

We define a differential equation or a stochastic integral as an appropriate limit to make sense of the expression:
Z T
f (s, ω)dBs (ω)
0

(n)
In the following, let tk := kT 2−n , k = {0, 1, · · · , 2n − 1}. Thus, we have Btnk := BkT 2−n .
We first note that one cannot define the above in the Riemann-Stieltjes sense (i.e., by partitioning the domain and taking
limits as the partition gets finer) for an arbitrary measurable f 1 . To gain further insight as to why this leads to an issue, we
discuss the following. Using the independent-increments property (that is (i) in Definition 11.1.1) of the Brownian motion
(e.g. via the construction of (11.2)), the following can be shown:

Lemma 11.2.1 In L2 (that is, mean-square) and hence in probability


X
lim (Btnk+1 − Btnk )2 = T.
n→∞
k

Observe the following [312].

Theorem 11.2.1 Define the total variation of the Brownian process in the interval [a, b] as:
X
T V (B, a, b) = sup |Btk+1 − Btk |
a≤t1 ≤t2 ≤···≤tk ≤b,k∈N
k

Almost surely, T V (B, a, b) = ∞.


P
Proof. By Lemma 11.2.1, and Theorem B.3.2, it follows that there exists some subsequence nm so that k (Btk+1
nm −
Btnk m )2 → b − a almost surely (see Theorem B.3.2). Now, if T V (B, a, b) < ∞, this would imply that
X X
2
m − Btnm ) ≤ sup |Btn
(Btnk+1 − Btnk | |Btnk+1 − Btnk | → 0,
k k+1
k
k k

as n → ∞, since by continuity of sample paths (which would then be uniformly continuous due to compactness of the
support, for any sample path with probability one) supk |Btnk+1 − Btnk | → 0. This would lead to a contradiction. ⋄
To appreciate some subtleties on stochastic integration, let us consider simple functions of the form:
n
2X −1
f (t, ω) = ek (ω)1{t∈[kT 2−n ,(k+1)T 2−n ]}
k=0

where n ∈ N. Let us define


1
However, this would be applicable if one has further regularities on the integrand f [339].
11.2 Stochastic Integration, the Itô Integral and Stochastic Differential Equations 249
Z T X
f (t, ω)dBt (ω) = ek (ω)[Btk+1 (ω) − Btk (ω)]
0 k
(n)
where tk = tk = kT 2−n , k = {0, 1, · · · , 2n − 1}. In the following, we use the notation: Btnk := BkT 2−n .
Now, note that if one has: X
f1 (t, ω) = B{kT 2−n } 1{t∈[kT 2−n ,(k+1)T 2−n ]}
k

it can be shown that Z T


E[ f1 (t, ω)dBt (ω)] = 0,
0
but instead with X
f2 (t, ω) = B{(k+1)T 2−n } 1{t∈[kT 2−n ,(k+1)T 2−n ]}
k

it can be shown that Z T


E[ f2 (t, ω)dBt (ω)] = T.
0

Thus, even though both f1 and f2 look to be reasonable approximations for some function f (t, ω), such as Bt (ω), the
integrals have drastically different meanings.
In particular the variations in the Bt process is too large to define an integration (in the usual sense of Riemann-Stieltjes),
RT
as we discussed above: It does make a difference on whether one defines 0 f (t, ω)dBt (ω) as an appropriate limit of a
sequence of expressions X
f (t∗j , ω)(Btnk+1 (ω) − Btnk (ω))
k

for different f (t∗j , ω) with (t∗j ∈ [tj , tj+1 ]. If we take t∗j = tj (the left end point), this is known as the Itô Integral. If we
take t∗j = 21 (tj + tj+1 ), this is known as the Stratonovich integral (and is denoted with f ◦ dBt , to distinguish it from the
R

Itô integral).

11.2.2 The Itô Integral

Itô’s integral will be well-defined2 , if we restrict the integrand f (t, ω) to be such that f (t, ω) is measurable on the σ-field
generated by {Bs , s ≤ t}. We define Ft to be the σ-algebra generated by Bs , s ≤ t. In other words, Ft is the smallest
σ-algebra containing sets of the form:

{ω : Bt1 (ω) ∈ A1 , · · · , Btk (ω) ∈ Ak }, tk ≤ t,

for Borel A1 , · · · , Ak . We also assume that all sets of measure zero are included in Ft (this operation is known as the
completion of a σ-field).

Definition 11.2.1 Let Nt , t ≥ 0, be an increasing family of σ-algebras of subsets of Ω. A process g(t, ω) is called Nt
adapted if for each t, g(t, ·) is Nt -measurable.

Definition 11.2.2 Let V(S, T ) be the class of functions:

f (t, ω) : [0, ∞) × Ω → R
2
In the theory of integration, as a student has seen over many courses, one chooses a definition of integration and identifies conditions
under which integration is possible. Of course, one expects that different integration concepts should be compatible whenever they are
simultaneously applicable. We have seen the Riemann integration and the Lebesgue integration, and how they both are defined as limits
of particular constructions. We will see in the following that the Itô integration has a similar flavour but with a very different construction.
This approach carries over to other types of integrations, such as the rough integral [135, 220].
250 11 Controlled Stochastic Differential Equations
RT
such that (i) f (t, ω) is B([0, ∞)) × F-measurable, (ii) f (t, ω) is Ft -adapted and (iii) E[ S f 2 (t, ω)dt] < ∞.

We will often take S = 0 in the following. For functions in V, the Itô integral is defined as follows: A function ϕ is called
elementary if it has the form: X
ϕ(t, ω) = ek (ω)1{t∈[tk ,tk+1 )}
k

with ek being Ftk -measurable. For elementary functions, we define the Itô integral as:
Z T X  
ϕ(t, ω)dBt (ω) = ek (ω) Btk+1 (ω) − Btk (ω) (11.3)
0 k

With this definition, it follows that for a bounded and elementary ϕ,


 Z T 2  Z T
E ϕ(t, ω)dBt (ω) = E[ ϕ2 (t, ω)dt]. (11.4)
0 0

This property is known as the Itô isometry. The proof follows from expanding the summation in (11.3) and using the
properties of the Brownian motion. Now, the remaining steps to define the Itô integral are as follows:
– Step 1: Let g ∈ V and g(·, ω) be continuous for each ω. Then, there exist elementary functions ϕn ∈ V such that
Z T
E[ (g − ϕn )2 dt] → 0,
0

as n → ∞. The proof here follows from the dominated convergence theorem.


– Step 2: Let h ∈ V be bounded. Then there exist gn ∈ V such that gn (·, ω) is continuous for all ω and n and
Z T
E[ (h − gn )2 dt] → 0.
0

One can follow Lusin’s theorem (Theorem D.5.1) to establish this result.
– Step 3: Let f ∈ V. Then, there exists a sequence hn ∈ V such that hn is bounded for each n and
Z T
E[ (f − hn )2 dt] → 0.
0

Here, we use truncation and then the dominated convergence theorem.

Definition 11.2.3 (The Itô Integral) Let f ∈ V(S, T ) and ϕn be an approximating sequence of elementary functions as
above given in Steps 1-2-3. The Itô integral of f is defined by
Z Z T
f (t, ω)dBt (ω) = lim ϕn (t, ω)dBt (ω),
n→∞ 0

where the convergence to the limit is in L2 (P ) in the sense; that is,


 Z T Z 2 
lim E ϕn (t, ω)dBt (ω) − f (t, ω)dBt (ω) =0
n→∞ 0

The existence of a limit is established through the construction of a Cauchy sequence and the completeness of L2 (P ),
the space of measurable functions with a finite second moment under P , with the corresponding norm. A computationally
useful result is the following (generalizing (11.4)).
11.2 Stochastic Integration, the Itô Integral and Stochastic Differential Equations 251

Corollary 11.2.1 For all f ∈ V(0, T )


 Z T 2  Z T 
E f (t, ω)dBt (ω) =E f 2 (t, ω)dt .
0 0

And thus, if f, fn ∈ V(0, T ) and


Z T
E[ (fn − f )2 dt] → 0,
0

then in L2 (P )
Z T Z T
fn (t, ω)dBt (ω) → f (t, ω)dBt (ω)
0 0

Example 11.3. Let us show that


Z t
1 2 1
Bs dBs = B − t. (11.5)
0 2 t 2
Rt
With Bjn := Btnj , define the elementary function: ϕn (ω) = Bjn (ω)1{t∈[tnj ,tnj+1 )} , it follows that E[ 0 (ϕn − Bs )2 ds] →
P

0. Therefore, the limit of the integrals of ϕn (ω), that is the L2 (P ) limit of j Bj (Bj+1 − Bjn ), will be the integral.
P n n
Observe now that
n
−(Bj+1 − Bjn )2 = 2Bjn (Bj+1
n
− Bjn ) + Bjn 2 − (Bj+1n
)2
and thus summing over j, we obtain
X X
−(Bj+1n
− Bjn )2 = 2Bjn (Bj+1
n
− Bjn ) + Bjn 2 − (Bj+1
n
)2 ,
j j

leading to X X
Bt2 − n
(Bj+1 − Bjn )2 = 2Bjn (Bj+1
n
− Bjn ) + B02 ,
j j

with B0 = 0. Now, taking the intervals [j, j + 1] arbitrarily small, we see that the first term converges to Bt2 −t (see Lemma
Rt
11.2.1) and the term on the right hand side converges to 2 0 Bs dBs , leading to the desired result. We will derive the same
result using Itô’s formula shortly. The message of this example is to highlight the computational method: Find a sequence
of elementary function which converges in L2 (P ) to f , and then compute the integrals, and take the limit as the intervals
shrink.

Remark 11.4. An important extension of the Itô integral is to a setup where ft is Ht -measurable, where Ht ⊂ Ft =
σ(Bs , s ≤ t). In applications, this is important to let us apply the integration to settings where the process that is integrated
is measurable only on a subset of the filtration generated by the Brownian process. This allows one to define multi-
dimensional Itô integrals as well. This is particularly useful for controlled stochastic differential equations, where the
control policies are measurable with respect to a filtration that does not contain that generated by the Brownian motion, but
the controller policy cannot depend on the future realizations of the Brownian motion either.

Remark 11.5 (Itô vs. Stratonovich Integrations). A curious reader may question the selection of choosing the Itô integral
over any other, and in particular the Stratonovich integral: Different applications are more suitable for either interpreta-
tion. In stochastic control, measurability aspects (of admissible controls) are the most crucial ones. If one appropriately
defines the functional or stochastic dependence between a function to be integrated or a noise process, the application of
either will come naturally: If the functions are to not look at the future, then Itô’s formula is appropriate. However, for
many applications involving white noise-like disturbances, physical processes, or stochastic stability, ergodicity [17, 193]
and smoothness properties of densities of solutions [154] to stochastic differential equations where one would build on
connections with geometric control theory [297] with piece-wise constant control action sequences replacing the driving
noise process, Stratonovich integral has been shown to be more relevant. Additionally, the Stratonovich integration has
desirable robustness properties with regard to the approximation of the Brownian noise, as will be discussed in Section
252 11 Controlled Stochastic Differential Equations

11.8. A conclusion is that the application itself should determine the right notion of the stochastic integral to be used, in
view of the assumptions imposed by the application.

11.2.3 The Itô Formula

Now that we have defined integration, we will study a generalization of the chain rule in classical calculus: Itô’s formula
allows us to take integrations of functions of processes.

Definition 11.2.4 We say v[0, T ) ∈ WH if


v(t, ω) : [0, T ) × Ω → R
is such that (i) v(t, ω) is B([0, ∞)) × F-measurable, (ii) v(t, ω) is Ht -adapted where Ht is as in Remark 11.4 and (iii)
RT
P ( 0 f 2 (t, ω)dt < ∞) = 1.

Definition 11.2.5 (Itô Process) Let Bt be a one-dimensional Brownian motion on (Ω, F, P ). A (one-dimensional) Itô
process is a stochastic process Xt on (Ω, F, P ) of the form
Z t Z t
Xt = X0 + b(s, ω)ds + v(s, ω)dBs (11.6)
0 0
Rt
where v ∈ WH so that v is Ht -adapted and P ( 0
v 2 (t, ω)dt < ∞) = 1 for all t ≥ 0. Likewise, b is also Ht -adapted and
Rt
P ( 0 b2 (t, ω)dt < ∞) = 1 for all t ≥ 0.

Instead of the integral form in (11.6), we may use the differential form notation:

dXt = bdt + vdBt ,

with the understanding that this means the integral form.

Theorem 11.2.2 [Itô Formula] Let Xt be an Itô process given by

dXt = bdt + vdBt .

Let g(t, x) ∈ C 1,2 ([0, ∞) × R) (that is, g is continuously differentiable, C 1 , in t, and twice continuously differentiable,
C 2 , in x). Then,
Yt = g(t, Xt ),
is again an Itô process and

∂g ∂g 1 ∂2g
dYt = (t, Xt )dt + (t, Xt )dXt + (t, Xt )(dXt )2
∂t ∂x 2 ∂x2
where
(dXt )2 = (dXt )(dXt )
with dtdt = dtdBt = dBt dt = 0 and dBt dBt = dt

Remark 11.6. Let us note that if instead of dBt , we only had a differentiable function mt so that dXt = bdt + vdmt , the
regular chain rule would lead to:

∂g ∂g
dYt = (t, Xt )dt + (t, Xt )(udt + vdmt ).
∂t ∂x
Note then that Itô’s Formula is a generalization of the ordinary chain rule for derivatives. The difference is the presence of
the quadratic term that appears in the formula; see also the discussion at the beginning of the chapter.
11.2 Stochastic Integration, the Itô Integral and Stochastic Differential Equations 253

Example 11.7. Let us compute Z t


Bs dBs .
0

View the above as an application of Itô formula with Xt = Bt so that dXt = dBt , and Yt = g(t, Xt ) = 21 Xt2 . Then, by
Itô’s formula,
1 1
dYt = d(g(t, Xt )) = Xt dXt + dt = Bt dBt + dt
2 2
and thus Z Z
1 1
dYs = Bt2 = Bs dBs + t.
2 2
R
Noting that dYs = Yt − Y0 , this result is in agreement with (11.5).

Itô’s Formula can be extended to higher dimensions by considering each coordinate separately.

11.2.4 Stochastic Differential Equations

Consider now an equation of the form:

dXt = b(t, Xt )dt + σ(t, Xt )dBt (11.7)

with the interpretation that this means


Z t Z t
Xt = X0 + b(s, Xs )ds + σ(s, Xs )dBs
0 0

Three natural questions are as follows: (i) Does there exist a solution to this differential equation? (ii) Is the solution
unique? (iii) How can one compute the solution?

Theorem 11.2.3 (Existence and Uniqueness Theorem) Let T > 0 and b : [0, T ] × Rn → Rn , σ : [0, T ] × Rn → Rn×m
be measurable functions satisfying:

|b(t, x)| + ∥σ(t, x)∥ ≤ C(1 + |x|), t ∈ [0, T ], x ∈ Rn

for some C ∈ R with ∥σ∥2 := Trace(σσ T ), and

|b(t, x) − b(t, y)| + |σ(t, x) − σ(t, y)| ≤ D(|x − y|), t ∈ [0, T ], x, y ∈ Rn

for some constant D. Let X0 = Z be a random variable which is independent of Bs , s ≥ 0 with E[|Z|2 ] < ∞. Then,
the stochastic differential equation (11.7) has a unique solution Xt (ω) that is continuous in t with the property that Xt is
RT
adapted to the filtration generated by {Z, Bs , s ≤ t} and E[ 0 |Xt |2 ] < ∞.

Proof Sketch. The proof of existence follows from a similar construction for the existence of solutions to ordinary differ-
ential equations: One defines a sequence of iterations:
Z t Z t
Ytk+1 = X0 + b(s, Ytk )ds + σ(s, Ytk )dBs
0 0

with Yt0 := X0 for all t ∈ [0, T ]. Then, the goal is to obtain a bound on the L2 -errors so that

lim E[|Ytm − Ytn |2 ] → 0,


m,n→∞

so that Ytn is a Cauchy sequence under the L2 (P ) norm; this is where the Lipschitz bounds in the hypothesis are utilized.
Call the limit X. The next step is to ensure that X indeed satisfies the equation and that there can only be one solution.
254 11 Controlled Stochastic Differential Equations

Finally, one proves that Xt can be taken to be continuous. ⋄

Let us appreciate some of the conditions stated above in the context of deterministic models.

Remark 11.8. Consider the following deterministic differential equations:



dx x
=4
dt t
with x(1) = 1 does not admit a unique solution on the interval [−1, 1].
– The differential equation
dx
= x2
dt
1
with x(0) = 1 admits the solution xt = 1−t and as t ↑ 1, the solution blows up in finite time so that there is no
solution for t ≥ 1.

The solution discussed above is what is called a strong solution. Such a solution is such that Xt is unique for a given
sample path. Furthermore, the solution is measurable on the filtration generated by the Brownian motion and the initial
variable (which can be seen by the construction of the integral, where each pre-limit approximation is measurable on
the filtration, and since L2 -limit implies a pointwise almost sure limit along a subsequence, the limit is also measurable,
assuming completeness of the filtration). Such a solution concept has an important engineering/control appeal in that the
solution is completely specified once the realizations of the Brownian motion (together with the initial state) are specified.
Weak solutions. In many applications, however, the conditions of Theorem 11.2.3 do not hold. In this case, one cannot
always find a strong solution. However, in this case, one may be able to find a solution which satisfies the probabilistic
flow in the system so that the evolution of the probabilities are well-defined: Note, however that, this solution may no
longer be adapted to the filtration generated by the actual Brownian motion and the initial state; but may be adapted to
some other Brownian process defined on some probability space. Such a solution is called a weak solution or a martingale
solution. While such a definition has a physical interpretation limitation in the sense that the input-output relation does
not correspond to one where the noise is an input and the solution is the output, this concept is instrumental in studying
controlled stochastic differential equations as we will discuss later in the chapter and is appropriate if one is concerned
with expected behaviour of the solutions. This concept is also related to the solution to the Fokker-Planck equation that we
will discuss further in the chapter in Section 11.2.6. For weak solutions, it suffices to have b to be bounded and only σ to
satisfy the Lipschitz continuity and the growth conditions provided that σ(·)σ T (·) has its eigenvalues uniformly bounded
from below (at least locally). We will discuss this further in the context of Girsanov’s measure transformation.

11.2.5 Some Properties of SDEs

Definition 11.2.6 A diffusion (also called Itô diffusion) is a stochastic process Xt (ω) satisfying a stochastic differential
equation of the form:
dXt = b(Xt )dt + σ(Xt )dBt , t ≥ s, Xs = x
where Bt is m-dimensional Brownian motion and b, σ satisfy the conditions of Theorem 11.2.3 so that

|b(x) − b(y)| + |σ(x) − σ(y)| ≤ D|x − y|.

Note that here b, σ only depend on x and not on t. Thus, the process here is time-homogenous.

Theorem 11.2.4 Let Xt be a diffusion and f be bounded and (Borel) measurable. Then, for t, h ≥ 0:

Ex [f (Xt+h )|Ft ](ω) = EXt (ω) [f (Xh )]


11.2 Stochastic Integration, the Itô Integral and Stochastic Differential Equations 255

Theorem 11.2.5 (Strong Markov Property) Let f be bounded and Borel, and τ be a stopping time with respect to Ft =
σ({Bs , s ≤ t}). Then, for h ≥ 0, conditioned on the event that τ < ∞:

Ex [f (Xτ +h )|Fτ ](ω) = EXτ (ω) [f (Xh )]

Definition 11.2.7 Let Xt be a time-homogenous Itô diffusion in Rn . The infinitesimal generator A of Xt is defined by:

Ex [f (Xt )] − f (x)
Af (x) = lim , x ∈ Rn ,
t→0 t
whenever f is so that the limit is defined.

Lemma 11.2.2 Let Yt = Ytx be an Itô process in Rn of the form:


Z t Z t
Ytx (ω) =x+ u(s, ω) + v(s, ω)dBs (ω).
0 0

Let f ∈ Cc2 (Rn ), that is f is twice continuously differentiable and has compact support, and τ be a stopping time with
respect to Ft with Ex [τ ] < ∞. Assume that u, v are bounded. Then,
Z τ X
∂2f
 
i ∂f 1X T
E[f (Yτ )] = f (x) + Ex u (s, ω) i (Ys ) + (vv )ij (s, ω) i j (Ys ) ds .
0 i
∂x 2 i,j ∂x ∂x

This lemma, combined with Definition 11.2.7 gives us the following result:

Theorem 11.2.6 Let dXt = b(Xt )dt + σ(Xt )dBt . If f ∈ Cc2 (Rn ), then,

∂2f
X 
i ∂f 1X T
Af (Xs ) = b (x) i (Xs ) + (σσ )ij (s, ω) i j (Xs )
i
∂x 2 i,j ∂x ∂x

A very useful result follows.

Theorem 11.2.7 (Dynkin’s Formula) Let f ∈ Cc2 (Rn ) and τ be a stopping time with Ex [τ ] < ∞. Then,
Z τ
Ex [f (Xτ )] = f (x) + Ex [ Af (Xs )ds]
0

Remark 11.9. The conditions for Dynkin’s Formula can be generalized. As in Theorem 4.1.5, if the stopping time τ is
bounded by a fixed constant, the conditions on f can be relaxed. Furthermore if τ is the exit time from a bounded set, then
it suffices that the function is C 2 (and does not necessarily have compact support).

Remark 11.10. [Martingale characterization of weak solutions] Consider a stochastic differential equation:dXt = b(Xt )dt+
σ(Xt )dBt . As we studied earlier, a probability measure P on the sample path space (or its stochastic realization Xt ) is
said to be a weak solution if under P
Z t
f (Xt ) − Af (Xs )ds (11.8)
0

is a martingale with respect to Mt = σ(Xs , s ≤ t), for any C 2 function f with bounded first and second order partial
derivatives. As noted earlier, every strong solution is a weak solution, but not every weak solution is a strong solution; every
such P admits a stochastic realization [180] but the stochastic realization may not be defined on the original probability
space as a measurable function of the original Brownian motion. For example, if Xt can be defined to be randomized, where
the randomization variables are independent noise processes, one could embed the noise terms into a larger filtration; this
will lead to a weak solution but not a strong solution since there is additional information required (that is not contained in
the original Brownian process).
256 11 Controlled Stochastic Differential Equations

11.2.6 Fokker-Planck equation

The discussion on the infinitesimal generator function (and Dynkin’s formula) suggests that one can compute the evolution
of the probability measure µt (·) = P (Xt ∈ ·), by considering for a sufficiently rich class of functions f ∈ D
Z
E[f (Xt )] = µt (dx)f (x).

Note that continuous and bounded functions are measure determining (as discussed in the proof of Theorem 10.8, see
[42, p. 13] or [119, Theorem 3.4.5]) and since smooth signals are dense among such functions, we can take f to be
smooth. Suppose that we assume that µt admits a density function and this is denoted by the same letter. Furthermore, let
p(x, t) := µt (x). By taking D to be the space of smooth signals with compact support, which is a dense subset of the space
of square integrable functions on R, using the expectation of the infinitesimal generator function equation (11.8), writing
Z t
1 ∂ 2 f (x) 2
 
d d d df
Z Z
µt (dx)f (x) = E[f (Xt )] = E[ Af (Xs )ds] = µt (dx) b(x) + σ (x)
dt dt dt 0 dx 2 ∂x2

and applying integration by parts (twice for the term on the right), we obtain that for a process of the form

dXt = b(Xt )dt + σ(Xt )dBt (11.9)

the following holds:

∂p(x, t) ∂ 1 ∂2 2
= − (b(x)p(x, t)) + (σ (x)p(x, t)) (11.10)
∂t ∂x 2 ∂x2
This is the celebrated Fokker-Planck equation. Notably, if there exists a stationary measure p, the time-independence on
the right hand side will lead to an ODE for this stationary measure.
The Fokker-Planck equation is a partial differential equation whose existence for a solution requires certain technical
conditions. As we discussed earlier, this is related to having a weak solution to a stochastic differential equation and in
fact they typically imply one another. Of course, the Fokker-Planck equation may admit a density as a solution, but it may
also admit a solution in a further weaker sense in that the evolution of the solution measure P (Xt ∈ ·) may not admit a
probability density function.

11.2.7 Rough Integration [Optional]

We end this section with a brief reflection on the limitations of the integrations noted above. From the way we have
constructed the Itô integral is that the integral is constructed as an L2 limit of approximations. In particular, (i) the integral
is not defined in a sample path sense (and typically only would allow for convergence in probability and thus almost
sure convergence along a subsequence, though this does occur also in a sample path almost sure sense under additional
conditions on the regularity of the integrand; see e.g. [339] or [312, p. 91]), and (ii) it is not continuous with respect to
the driving noise. A pathwise theory of solutions to differential equations would thus be a natural goal to arrive at; and
this is attained by what is known as rough integration. The fundamental insight of rough paths theory is that the issue
of defining solutions to differential equations driven by an irregular signal X = (X 1 , ..., X d ) can be reduced to defining
Rt
iterated integrals of the form s (X i (r) − X i (s))dX j (r). Rough paths theory allows for Hölder continuous driving signals
- or in the case of stochastic differential equations, stochastic processes that are almost surely Hölder. Recall the definition
of Hölder continuity: Define for α ∈ (0, 1] the space C α ([0, T ], Rd ) of α-Hölder functions f : [0, T ] → Rd equipped with
the norm ∥f ∥α := supt̸=s |f (t)−f (s)|
|t−s|α . Now, if X is a signal that is α-Hölder continuous with α ∈ (1/3, 1/2] and F is a
smooth function, then for a partition of [0, t], P = {0 = t0 < ... < tn = t} we have that in the integral
Z t n Z
X tk+1
F (X(r))dX(r) = F (X(r))dX(r)
0 k=0 tk
11.3 Controlled Stochastic Differential Equations and the Hamilton-Jacobi-Bellman Equation 257
n Z
X tk+1
= F (X(tk )) + F ′ (X(tk ))(X(r) − X(tk )) + O(|r − tk |2α )dX(r)
k=0 tk

Xn  Z tk+1
= F (X(tk ))(X(tk+1 ) − X(tk )) + F ′ (X(tk )) (X(r) − X(tk ))dX(r)
k=0 tk


+ O(|tk+1 − tk | ) , (11.11)

Rt
as 3α > 1, the remainder term should go to 0. This reduces the problem of defining the integral 0 F (X(r))dX(r) to
Rt
just defining tkk+1 (X(r) − X(tk ))dX(r). We take the right hand side as a definition of the left hand side, so long as we
define the iterated integral first. However, if X is irregular then the iterated integral does not exist as a Riemann-Stieltjes
integral and therefore must be defined. This leads to a new construction, called the rough integral. A rough path above
a signal X is a pair Xs,t = (Xs,t , Xs,t ) where Xs,t is the increment of X and Xs,t is a definition or postulation of
Rt
the iterated integral s (X(r) − X(s))dX(r) [136]. Then, one constructs the definition in a sense that it is compatible
Rt
with the usual integration notions. For example, if X ∈ C 1 and s (X(r) − X(s))dX(r) is the Riemann-Stieltjes integral,
Rt
Xs,t := (X(t)−X(s), s (X(r)−X(s))dX(r)) is consistent with such a postulation. The main utility of rough integration
is that, by imposing conditions on the rough integration, solutions to equations of the form [136, Theorem 4.10]

dY 1 = b1 (Y 1 )dt + σ(Y 1 )dX1 ,

is continuous in the driving noise. It should be noted that the conditions impose that the rough integral definition itself is
continuous in the driving noise.

11.3 Controlled Stochastic Differential Equations and the Hamilton-Jacobi-Bellman Equation

11.3.1 Revisiting the deterministic optimal control problem in continuous-time

Consider
dx
= f (x, u), x(0) = x
dt
and suppose that the goal is to minimize
Z T
J(γ) = c(s, x(s), u(s))ds + P (x(T ))
t0

over all feedback control policies γ, where we take c to be continuous and bounded. Via Bellman’s principle, as in Theorem
5.1.3, we define value functions:
Z T
V (t, x) = inf c(s, x(s), u(s))ds + P (x(T ))
γ t

with the terminal condition


V (T, x(T )) = P (x(T ))

Remark 11.11. For the existence of an optimal policy, very mild conditions can be arrived at via the theory of Young
measures: See Section 11.5.1 for a detailed analysis leading to general existence conditions.

In the following, we first present an informal and non-rigorous derivation for an optimality equation, but the optimality
analysis will be rigorously justified in Theorem 11.3.1. Applying Bellman’s principle from the theory studied earlier, for a
policy to be optimal it looks reasonable to arrive at the following (which will be justified shortly):
258 11 Controlled Stochastic Differential Equations
Z t+h
V (t, x) = inf c(s, x(s), u(s))ds + V (t + h, x(t + h)) (11.12)
γ t

or
 Z t+h 
0= inf c(s, x(s), u(s))ds + V (t + h, x(t + h)) − V (t, x) (11.13)
γ t

for all h ≥ 0. Consider then:


 
R t+h
inf γ t
c(s, x(s), u(s))ds + V (t + h, x(t + h)) − V (t, x)
0 = lim
h→0 h
Now, for h small, we have that x(t + h) = x(t) + f (x(t), u(t))h + o(h), where o(h)/h → 0 as h → 0. If we assume that
V is continuously differentiable in its entries, we then have

V (t + h, x(t + h)) = V (t, x(t)) + Vt (x, t)h + (Vx (t, x) · f (x, u))h + o(h),

leading to  
R t+h
t
c(s, x(s), u(s))ds + Vt (t, x)h + (Vx (t, x) · f (x, u))h + o(h)
0 = lim inf
h→0 γ h
Assuming that o(h)/h → 0 uniformly for all control policies and that c(s, x(s), u(s)) is continuous in s (which is clearly
not justified for an arbitrary policy!), we arrive at
 
0 = inf c(t, x(t), u(t)) + Vt (t, x) + (Vx (t, x) · f (x, u))
γ

In a more standard form, this leads to, provided the minimum exists,
 
−Vt (t, x) = min c(t, x, u(t)) + (Vx (t, x) · f (x, u(t))) (11.14)
u(t)

This is the celebrated HJB (Hamilton-Jacobi-Bellman) equation. This defines a partial differential equation with boundary
condition V (T, x) = P (x) or V (T, x(T )) = P (x(T )).
The above analysis has several gaps: we imposed the value functions to be so that local linearized approximations could
be made and some uniformity assumptions were not even justified. Nonetheless, as is often the case in applied mathemat-
ics, heuristic reasoning may lead to important equations whose validity however then needs to be rigorously justified. In
particular, the above leads to an important equation which is a surprisingly strong result, as established in the following
verification theorem:

Theorem 11.3.1 [Optimality of HJB Solutions] Let V (t, x) be C 1 (i.e., continuously differentiable) in both t and x, and
solve the HJB equation (11.14). Suppose further that the policy γ satisfies (11.14) with u(t) = γ(t). Then, γ is optimal.

Proof. Let V (t, x) be C 1 in both entries. Consider any admissible policy γ, which (under any history dependent measurable
policy) can be viewed to be a function of time without any loss since the problem is deterministic. Let the HJB equation
hold:  
∂V ∂V
(t, x) + min (t, x) · f (x, u) + c(t, x, u) = 0, V (T, x) = P (xT ),
∂t u ∂x
and thus, for any policy with ut = γ(t):
 
∂V ∂V
(t, x) + (t, x) · f (x, γ(t)) + c(t, x, γ(t)) ≥ 0, V (T, x) = P (xT ),
∂t ∂x
11.3 Controlled Stochastic Differential Equations and the Hamilton-Jacobi-Bellman Equation 259

and thus, for any policy γ (which is open-loop without any loss in optimality for deterministic systems)
 
∂V ∂V
(t, x) + (t, x) · f (x, γ(t)) + c(t, x, γ(t)) ≥ 0.
∂t ∂x

Now, consider V (t, xγt ) where γ denotes the explicit dependence on the policy. We have that

dV (t, xγt ) ∂V ∂V
= (t, xγt ) + (t, x) · f (x, u)|x=xγt ,u=γ(t)
dt ∂t ∂x
By the above, we have then

∂V (t, xγt ) ∂V
−( + (t, x) · f (x, u)|x=xγt ,u=γ(t) ) ≤ c(x, γ(t))
∂t ∂x
and
dV (t, xγt )
− ≤ c(x, γ(t)) (11.15)
dt
Taking the integral and noting that this holds for any γ, we arrive at
Z T Z T
V (0, x0 ) ≤ c(xt , γ(t))dt + V (T, xγT ) = c(xt , γ(t))dt + P (xγT )
0 0

Note that the initial value is independent of control and hence V (0, x0 ) is a lower bound for any control. Equality holds if
the HJB is satisfied by some admissible control policy, which would then be optimal. ⋄

Remark 11.12. For some generalizations on HJB and optimal control:


(i) One can relax the regularity conditions on V so that V may not be differentiable everywhere (leading to solution
concepts such as viscosity solutions). An intuitive way to appreciate viscosity solutions is to consider the verification
theorem above and replace V at a neighborhood of a point x where V is not differentiable with a continuously
differentiable function ϕ which satisfies two properties: With V (t, x) = ϕ(t, x), if V (s, y) − ϕ(s, y) has its local
minimum at (t, x), then a version of (11.14) holds with
 
−ϕt (t, x) ≤ min c(t, x, u(t)) + (ϕx (t, x) · f (x, u(t))) ,
u(t)

which then leads to ϕ(t, x) − ϕ(s, y) ≥ V (t, x) − V (s, y) for (s, y) in a neighborhood of (t, x). Noting (ϕx (t, x) ·
f (x, u(t)) as the partial derivative of ϕ(t, ·), and moving it to the left hand side, by taking s = t + h, y = xt+h , we
arrive at (11.15) and this leads to the lower bound property of V . For the other direction, if V (s, y) − ϕ(s, y) has
its local maximum at (t, x), we have that ϕ(t, x) − ϕ(t, y) ≤ V (t, x) − V (t, y) in a neighborhood of (t, x) and in
the equation  
−ϕt (t, x) ≥ min c(t, x, u(t)) + (ϕx (t, x) · f (x, u(t))) ,
u(t)

fix the control policy attaining the minimum, and then arrive that
Z T
V (0, x0 ) ≥ c(xt , γ(t))dt + P (xγT )
0

Note the parallels in the argumentation, in terms of lower bounds and the attainability, with that in Theorem 5.5.3
in the order of the inequalities, as well as with the ACOI in Theorem 7.1.3.
(ii) Instead of sufficiency, one can arrive at necessary conditions via what is known as the maximum principle via
variational local optimality conditions. We refer the reader to [217] for a rather comprehensive and accessible
discussion and [130] for the stochastic setup.
260 11 Controlled Stochastic Differential Equations

11.3.2 The stochastic case and classes of admissible policies

Suppose now that we have a controlled system:

dXt = b(t, Xt , ut )dt + σ(t, Xt )dBt ,

where ut ∈ U is the control action variable. We assume that ut is measurable at least on Ft (but we can restrict this further
so that it is measurable on a strictly smaller sigma field). Thus, the differential equation is well defined as an Itô integral.
We will assume that a solution exists.
Often, one has a time-homogenous diffusion model, which is more suitable for infinite horizon analysis.

dXt = b(Xt , Ut )dt + σ(Xt )dBt (11.16)

Control policies and existence of solutions

As we discussed extensively throughout the notes, the selection of the control actions need to be measurable with respect
to some information at the controller and this dependency leads to fundamental differences on the behaviour of solutions
and optimization methods.

(i) If for every t, ut is measurable on the filtration generated by Xt , then the policy is called admissible.
(ii) If the control at time t is only a function of Xt and t, then the policy is called a Markov policy. If it only depends on
Xt , then it is stationary. Randomization also is possible, but this requires a more careful treatment when compared
with the discrete-time counterpart [15].
(iii)One often considers adapted open-loop policies; these are policies which are measurable with respect to σ(X0 , (Bs , s ≤
t)) at time t.
(iv)One further relaxes these with non-anticipative policies; these are policies which satisfy the condition that
X0 , (Us , Bs ) is independent of Bt − Bs for every t > 0.

The above entail subtle distinctions on whether they lead to strong solutions: To ensure existence and uniqueness of strong
solutions, consider the following assumptions on the drift b and the diffusion matrix σ .
(A1) Lipschitz continuity: The functions σ, b are Lipschitz continuous in x (uniformly with respect to the other variables
for b). In other words, for some constant C > 0

|b(x, ζ) − b(y, ζ)|2 + ∥σ(x) − σ(y)∥2 ≤ C|x − y|2


p
for all x, y and ζ ∈ U, where ∥ · ∥is an appropriate metric on matrices such as ∥σ(x)∥ = Traceσ(x)σ T (x).
Furthermore, b is jointly continuous in (x, ζ).

(A2) Affine growth condition: b and σ satisfy a global growth condition of the form

sup |⟨b(x, ζ), x⟩| + ∥σ(x)∥2 ≤ C0 (1 + |x|2 ) ∀x


ζ∈U

for some constant C0 > 0.


Recall that in Theorem 11.2.3 we had noted that for a control-free stochastic differential equation, if b and σ satisfy
Lipschitz regularity and growth properties, then a strong solution exists. However, if we only restrict that the control policy
is measurable, with no additional assumptions, it is not guaranteed that a strong solution for Xt would exist. In view of this
note, we state the following on the existence of strong or weak solutions under various classes of control policies:
(i) If b and σ satisfy similar regularity conditions, as in Theorem 11.2.3 with uniformity over control actions, and
the control policies are non-anticipative, then the existence of strong solutions under the hypotheses above follows
11.3 Controlled Stochastic Differential Equations and the Hamilton-Jacobi-Bellman Equation 261

from a similar argument [15, Theorem 2.2.4] (by viewing the control as an exogenous process). We also note
that the above global Lipschitz constant C can be relaxed to a local one, that is one may have for all R > 0,
|b(x, ζ) − b(y, ζ)|2 + ∥σ(x) − σ(y)∥2 ≤ CR |x − y|2 for all ∥x∥, ∥y∥ ≤ R; here one considers solutions up to an
exit time, establishes a uniqueness result and takes take the exit set boundary to infinity.
(ii) If admissible (or feedback policies; that is, policies which are σ(X[0,t] )-measurable) are considered, one can
apply measure transformation (via Girsanov’s method), establish a strong solution for the control-free term
dXt = σ(Xt )dBt , and then construct a solution under the original space; which then leads to a weak solution
for the original space. Accordingly, under the hypothesis (A2) together with a Lipschitz condition as in (A1) but
only for σ, for any admissible control (11.16) has a unique weak solution [15, Theorem 2.2.11].
(iii)Furthermore, under hypotheses (A1)–(A2) and under any stationary Markov strategy there is a unique strong solu-
tion which is a strong Feller process [15, Theorem 2.2.12].
(iv)Finally, if σ were allowed to also depend on control in general, existence in this setup is a non-trivial problem, since
measure change arguments cannot be directly applied when the control is present in the diffusion term. However, if
non-anticipative policies are considered, then under strict Lipschitz and growth conditions, a strong solution exists.
We refer the reader to [15, Section 2] as well as [33, 94, 95] for a detailed analysis and literature review. Notably,
if we replace σ(x) by σ(x, ζ), if σ(·, v(·)) is Lipschitz continuous for stationary v then there is a unique strong
solution. But in general stationary policies are just measurable functions, and existence of suitable strong solutions
is more delicate (see, [15, Remarks 2.3.2]).
In the following, we will consider admissible policies and the solution concept will be taken to be of weak solutions so
that the expected cost criteria are well-defined and the limitations on having strong solutions are not presented apriori.
Nonetheless, we will observe that under verification theorems optimal policies (even among those which are admissible)
will be Markov and the issues on existence of strong solutions will not be as critical, see the discussion above.
The reader is encouraged to revisit the verification theorems for discrete-time problems: These are Theorem 5.1.3 for finite
horizon problems, Theorem 5.5.3 for discounted cost problems, and Theorem 7.1.1 for average cost problems. One can see
that the essential difference is to express the expectations through Dynkin’s formula and the differential generators.

Finite Horizon Cost Criterion

Suppose that given a control policy, the goal is to minimize


Z T
E[ c(s, Xs , us )ds + cT (XT )],
0

where cT is some terminal cost function.


As in Chapter 5, if a policy is optimal, we will arrive at the following equation for every possible state:
Z t
V (r, Xr ) = min E[ c(s, Xs , us )ds + V (t, Xt )|Xr ].
γ r

In the following, following the same flow of ideas as in the deterministic case in Section 11.3.1, we first provide a rather
informal derivation of the optimality equation, but the formal verification result will be precise. We assume that V (s, x) is
C 2 in x and C 1 in s. Then,
 Z t 
0 = min E[ c(s, Xs , us )ds + V (t, Xt )|Xr ] − V (r, Xr )
u r

In particular,
R r+h
E[ r c(s, Xs , us )ds + V (r + h, Xr+h )|Xr ] − V (r, Xr )
 
0 = lim min
h→0 u h
262 11 Controlled Stochastic Differential Equations

Now, if V is so that it is in the domain of the generator for every control policy, with
X ∂V 1X i ∂2V
Lut V (t, x) = bi (t, x, u) (t, x) + σ (t, x)σ j
(t, x) (t, x)
i
∂xi 2 i,j ∂xi ∂xj

applying the mean-value theorem, assuming it would hold for now, we arrive at
 
u ∂V
min c(s, x, u) + Ls V (s, x) + (s, x) = 0 (11.17)
u ∂s

Thus, if a policy is optimal, it needs to satisfy the above property provided that V satisfies the necessary regularity condi-
tions under the considered set of policies to validate the operations above and indeed the above would also be sufficient for
optimality by the analysis to follow. However, as in the deterministic case, the analysis above is informal and we have not
presented precise conditions under which the above would hold. As we have seen before in the earlier chapters, verification
theorems show that a policy that satisfies the verification is optimal over all admissible policies:

Theorem 11.3.2 (Verification Theorem) Consider: dXt = b(t, Xt , ut )dt + σ(t, Xt , ut )dBt . Suppose that V is C 2 in x
and C 1 in t is so that:
 
∂V u
(t, x) + min Lt V (t, x) + c(t, x, u) = 0, V (T, x) = cT (x), (11.18)
∂t u∈U

Then, an admissible control policy which achieves the minimum for every (t, x) is optimal.

Proof.
∂V
The equation ∂t (t, x) + minu {Lu V (t, x) + c(t, x, u)} = 0 implies that for any admissible control realization:


c(t, x, u) ≥ − V (t, x) − Lu V (t, x), (11.19)
∂t
and as in the deterministic case (in Theorem 11.3.1), for any admissible control policy γ
T Z T
−∂
Z
γ u u u γ
E [ V (s, Xs ) − Ls V (s, Xs )] ≤ E [ c(s, Xsu , us )].
0 ∂s 0

Using Itô’s rule,


T
−∂V
Z
E[V (0, X0 )] = E[ ( (s, Xsγ ) − Lγs V (s, Xsγ ))ds] + E[V (T, XTγ )],
0 ∂s
and thus, we obtain that for any admissible control
Z T
E[V (0, X0 )] ≤ E[ c(s, Xsγ , us )ds + cT (XTγ )].
0

On the other hand, a policy γ ∗ which satisfies the equality in (11.18), leads to an equality in the above and optimality. ⋄

Example 11.13. [Optimal portfolio selection] We consider a continuous-time version of a problem considered in Exercise
2.9.2: A common example in finance applications is the portfolio selection problem where a controller (investor) would
like to optimally allocate his wealth between a stochastic stock market and a market with a guaranteed income (see [312]):
Consider a stock with an average return µ > 0 and volatility σ > 0 and a bank account with interest rate r > 0. These are
modeled by:

dSt = µSt dt + σSt dBt , S0 = 1


dRt = rRt dt, R0 = 1 (11.20)
11.3 Controlled Stochastic Differential Equations and the Hamilton-Jacobi-Bellman Equation 263

Suppose that the investor can only use his own money to invest and let ut ∈ [0, 1] denote the proportion of the money that
he invests in the stock. This implies that at any given time, his wealth dynamics is given by:

dXt = µuXt dt + σuXt dBt + r(1 − ut )Xt dt,


 
or dXt = µu + r(1 − u) Xt dt + σuXt dBt . Suppose that the goal is to maximize E[log(XT )] for a fixed time T (or
minimize −E[log(XT )]). In this case, the Bellman equation writes as:
2
 
∂ ∂ 2 2 21 ∂
0 = V (t, x) + min (µu + r(1 − u))x V (t, x) + σ u x V (t, x) ,
∂t u ∂x 2 ∂x2

with V (T, x) = − log(x). With a guess of the value function of the form V (t, x) = − log(x) + bt , one obtains an
ordinary differential equation for bt with terminal condition bT = 0. It follows that, if µ−r
σ 2 ∈ [0, 1] the optimal control is
ut (x) = µ−r
σ2 , leading to V (t, x) = − log(x t ) − C(T − t)), for some constant C.

Example 11.14. [The Linear Quadratic Regulator and Continuous-Time Riccati Equation] Consider a continuous-time
counterpart of the LQG problem studied in Section 5.3: Let

dxt = Axt dt + But dt + DdBt ,

where xt , ut are all R-valued, with the goal of minimizing:


Z N
J(x, γ) = Exγ [( Qx2t + Ru2t ) + QN x2N ],
0

where R > 0, QN > 0, Q ≥ 0. By the HJB equation, we have

1 ∂2
 
∂ ∂ 2 2 2
0 = V (t, x) + min V (t, x)(Ax + Bu) + V (t, x)D + Qx + Ru
∂t u∈R ∂x 2 ∂x2

Taking V (t, x) = Pt x2 + Kt with KN = 0, PN = QN , we then have


   
−Pt′ x2 − Kt′ = min 2Pt x(Ax + Bu) + Pt D2 + Qx2 + Ru2 = min 2APt x2 + 2Pt Bxu + Ru2 + Pt D2 + Qx2
u∈R u∈R


By completion of squares: 2Pt Bxu + Ru2 = ( Ru + √1 Pt Bx)2 − 1 2 2 2
R R Pt B x , we have that the optimal control is

1
ut = − Pt Bxt
R
where
1 2 2 2
−Pt′ x2 − Kt′ = (2APt + Q − P B )x + Pt D2 .
R t
Thus, we arrive at (the continuous-time Riccati equation)
1 2 2
−Pt′ = 2APt + Q − P B , PN = QN
R t
and
Kt′ = −Pt D2 , KN = 0.

11.3.3 Discounted Infinite Horizon Cost Criterion

Suppose that given a control policy, the goal is to minimize


264 11 Controlled Stochastic Differential Equations
Z
E [ e−λs c(Xs , us )ds].
γ
0

In this section, we will consider a time-homogenous setup

dXt = b(Xt , ut )dt + σ(Xt , ut )dBt , (11.21)

and let us define


X ∂g 1X i ∂2g
Lu g(x) = bi (x, u) (x) + σ (x, u)σ j
(x, u) (x) (11.22)
i
∂xi 2 i,j ∂xi ∂xj

In this case, we have the following result:

Theorem 11.3.3 (Verification Theorem) Suppose that V is C 2 and limt→∞ e−λt Exγ [V (Xtu )] = 0 under any admissible
policy γ and x. Let
 
u
min L V (x) − λV (x) + c(x, u) = 0. (11.23)
u∈U

Then, an admissible control policy which achieves the minimum above for every x is optimal.

Proof. Under any admissible policy γ and its control realization ut at time t, we have that Lγ V (xt )−λV (xt )+c(xt , ut ) ≥
0. This then leads to
e−λt c(xt , ut ) ≥ e−λt (−Lγ V (xt ) + λV (xt ))
Using Itô’s rule for V (Xt )e−λt one obtains:
Z t Z t
E γ [V (X0 )] − e−λt E[V (Xtγ )] = E γ [ e−λs (−Lγ V (Xs ) + λV (Xs ))ds] ≤ E γ [ e−λs c(xs , us )ds]
0 0

Taking t → ∞, we show that V (0) is a lower bound. Proceeding as before in the proof of Theorem 11.3.2, under equality,
we have that the lower bound is attained. ⋄
The above is a sufficiency analysis. One could say more with regard to necessity as well [15, Theorem 3.5.6] via an
analysis based on the theory of partial differential equations: In addition to mild growth conditions [15, Section 2.2], if (i)
c : Rd × U → R+ is continuous and locally Lipschitz in x uniform in u, and (ii) if σ(x), not depending on u, is locally
uniformly elliptic, i.e., σ(x)σ T (x) has its eigenvalues locally bounded from below, then (11.23) admits a unique solution,
which is bounded, and which serves as the solution to the optimal cost. Thus, one has a complete characterization of
optimality under additional regularity conditions. These also carry over to the average-cost setup presented in the following.

11.3.4 Average-Cost Infinite Horizon Cost Criterion

Suppose that given a control policy, the goal is to minimize


Z T
1 γ
lim sup E [ c(Xs , us )ds].
T →∞ T 0

Once again here we consider a time-homogenous setup with the controlled equation (11.21) and generator (11.22).

Theorem 11.3.4 (Verification Theorem) Suppose that V is C 2 and η ∈ R so that


 
min Lu V (x) − η + c(x, u) = 0, (11.24)
u∈U
11.4 Partially Observed Case, Girsanov’s Theorem and Separated Policies 265

and
Exγ [V (X0 ) − V (XTγ )]
lim sup = 0, (11.25)
T →∞ T

under every admissible policy γ and initial state x. Then, an admissible control policy which achieves the minimum for
every x is optimal.

Proof. Once again for any admissible control policy γ, Lγ V (xγt )−η +c(xγt , ut ) ≥ 0, leading to c(xγt , ut ) ≥ η −Lγ V (xγt ).
Then, one obtains Z T Z T
c(xγt , ut )dt ≥ ηT − Lγ V (xγt dt.
0 0
Dividing by T and taking the limit superior, via (11.25), we have the lower bound. The achievability holds under equality
(11.26). ⋄

The convex analytic method

The analysis we made in Chapter 5 and 7 applies to the diffusion setting as well. In particular, a discounted HJB equation
plays the role of the discounted cost optimality equation. For the average cost problems, one can apply either a vanishing
discount approach or an convex-analytic approach. We refer the reader to [15], [61, 66] and [295].

11.3.5 Control up to an Exit Time

In some applications, one studies cost criteria of the type:


Z τ
γ
E [ c(Xs , us )ds + h(Xτ )],
0

/ S} for some S ∈ Rn which is a set with a smooth boundary ∂(S) and h is a terminal cost
where τ = inf{t ≥ 0 : Xt ∈
function.
We again consider the time-homogenous setup with controlled equation (11.21) and generator (11.22).

Theorem 11.3.5 (Verification Theorem) Suppose that V is C 2 and η ∈ R so that


 
u
min L V (x) + c(x, u) = 0, x ∈ S; V (x) = h(x) on ∂(S). (11.26)
u∈U

Then, an admissible control policy which achieves the minimum for every x is optimal.

Remark 11.15. All the equations stated in the verification theorems noted above demonstrate the strict connections with
the theory of partial differential equations. Indeed, existence and optimality results can be utilized to obtain direct and very
strong results, see [15, Chapter 3]. The regularity conditions on the value function can also be relaxed.

11.4 Partially Observed Case, Girsanov’s Theorem and Separated Policies

Consider a partially observed setup with


Z
Yt = h(Xs )ds + Bt (11.27)
266 11 Controlled Stochastic Differential Equations

for some independent Brownian process Bt .


Such a setup leads to a number of technical difficulties. The analysis (especially for the case with measurements that
are not linear and Gaussian) can be quite subtle due to the fact that the control policy (only restricted to be measurable
in general) may lead to issues on the existence of strong solutions for a given stochastic differential equation since the
control policy may couple the state dynamics with the past in an arbitrarily complicated, though measurable, way and
hence violating the existence conditions for strong solutions to stochastic differential equations. Even for linear models,
the analysis requires some careful reflection: Lindquist [218] provides a detailed account on this aspect and provides a
general separation theorem provided that the control laws are among those which lead to the existence of a solution to the
controlled stochastic differential system, generalizing e.g. the analysis in [205] (where control laws of the Lipschitz type
in the conditional estimate are considered) and [334] (where Lipschitz control policies in y[0,t] are considered) to ensure
the existence of strong solutions.
To avoid such technical issues on strong solutions, relaxed solution concepts were introduced and studied in the literature
based on measure transformation due to Girsanov [33, 94, 95] (see Exercise 11.10.2 for a heuristic derivation). Now,
consider the measurement model given in (11.27).
For a moment, suppose that h ≡ 0, that is Yt is just an independent process. Suppose also that there is no control in the
diffusion process dXt = b(Xt )dt+σ(Xt )dBt′ . In this case, it is evident that the measurement process and the noise process
are independent and let us call this probability measure on the processes as Q := PX × QY . In this case, consider for any
measurable bounded f on the paths of X[0,t] and Y[0,t] as:
Z
EQ [f (X[0,t] , Y[0,t] )|Y[0,t] ] = f (x[0,t] , Y[0,t] )PX (dx[0,t] ),

since the measurement processes Yt gives no information on the state process Xt . Thus, the computation is quite simple in
this case.
Now, consider our original process given by (11.27) where h is non-zero. Let P be the joint probability measure on the
state and the measurement processes. Since P ≪ Q, we have that
Z Z
f (x[0,t] , y[0,t] )P(dx[0,t] , dy[0,t] ) = G(x[0,t] , y[0,t] )f (x[0,t] , y[0,t] )Q(dx[0,t] , dy[0,t] )

for some Q-integrable function G, which is the Radon-Nikodym derivative of P with respect to Q. It turns out that under
mild conditions, we have that
dP
Rt Rt
h(xs )dys − 12 |h(xs )|2 ds
Gt := G(x[0,t] , y[0,t] ) = =e 0 0
dQ
Rt
with 0 h(xs )dys being a stochastic integral, this time with respect to the random measure/process Ys . This relation allows
us to view the partially observed problem as one with independent measurements, with the dependence pushed to the
Radon-Nikodym derivative, not unlike what was done in Section 10.4.2 (see also Exercise 11.10.2).
Therefore,
R
x[0,t] ,y[0,t]
f (x[0,t] , y[0,t] )G(x[0,t] , y[0,t] )Q(dx[0,t] , y[0,t] )
EP [f (X[0,t] , Y[0,t] )|y[0,t] ] = R (11.28)
x[0,t]
G(x[0,t] , y[0,t] )Q(dx[0,t] , y[0,t] )

This equation is known as the (Kushner-)Kallianpur-Striebel formula. If we focus on the numerator and focus on Xt only,
this is known as the unnormalized filter [203].

11.4.1 Non-linear filtering in continuous time and Zakai’s equation

We will now study the evolution of the numerator


R y above where we restrict f to be a function of the current state only. In
particular, we will study the evolution of µt [0,t] (dx)f (x) = EP [f (Xt ), y[0,t] ], where the notation [·, y[0,t] ] means that we
11.5 Existence of Optimal Policies under Full, Partial and Decentralized Information 267

restrict the measurements y[0,t] to be fixed (but we are not computing the conditional measure), as measurements are also
collected. Compare this with the Fokker-Planck equation (11.10) in which case measurements do not exist.
Now, under Q, we have that the X and the Y process are independent. Note now that we can write

EP [f (Xt ), y[0,t] ] = EQ [f (Xt )Gt ],

where Gt is as defined earlier in this section. This relation will make the analysis below relatively immediate, following
the analysis of the Fokker-Planck equation (11.10). We will follow a similar reasoning, except now we will also consider
the realizations of the measurements by considering Gt as a variable which adds a time dependence. We have then:

d d d
Z
y
µt [0,t] (dx)f (x) = EP [f (Xt ), y[0,t] ] = EQ [f (Xt )Gt ]
dt dt dt
and thus,
1 ∂ 2 f (x) 2
 
d df dGt
Z Z
y
µt [0,t] (dx)f (x) = µQ
t (dx) b(x)G t + σ (x)Gt + f
dt dx 2 ∂x2 dt
As before, applying integration by parts (twice for the term in the middle), and then computing via Itô’s formula the
derivative of dGdt = Gt (hdY ) (this can be obtained by writing Gt = e
t Zt
where Zt solves dZt = − 12 |h(xt )|2 dt +
h(xt )dBt ), and then writing
y
µt [0,t] (dx) = µQ
t (dx)Gt

we obtain that for a process of the form

dXt = b(Xt )dt + σ(Xt )dBt (11.29)

with measurements given in 11.27), the non-normalized filter density (provided a smooth one exists) py (x, t) evolves as:

1 ∂2 2
 
y ∂ y
dp (x, t) = − (b(x)p (x, t)) + (σ (x)p (x, t)) dt + py (x, t)hdy
y
(11.30)
∂x 2 ∂x2

If one normalizes the above, this time by conditioning on y (that is, by dividing the above with the expectation over the
random measurements by integrating over all state values under the measure P), one arrives at another important equation,
known as the Kushner-Stratonovich equation. In particular, for a given function f
R y
p (x, t)f (x)dx
E[f (Xt )|Y[0,t] = y[0,t] ] = R y
p (x, t)dx

11.5 Existence of Optimal Policies under Full, Partial and Decentralized Information

11.5.1 A related existence discussion in deterministic continuous-time

We first revisit a version of the deterministic optimal control problem considered in Section 11.3.1. It is instructive to
discuss here various control topologies that are already well-known in classical control theory (when there is a single
controller who has access to the state variable).
In the following, we build on [271]. In deterministic nonlinear, geometric, and continuous-time control, properties on
stabilizability, controllability, and reachability are drastically impacted by the restrictions on the classes of allowed controls
(e.g., continuous, Lipschitz, finitely differentiable, or smooth control functions in the state or time when control is open-
loop [71, 179, 267, 291]) and naturally the control topology induced is dictated by the class of admissible controls. For
optimal control, to allow for continuity/compactness arguments, apriori imposing compactness over spaces of measurable
functions would be an artificial restriction, and the use of powerful theorems such as the Arzela-Ascoli theorem which
necessarily entail (usually very restrictive and suboptimal) conditions on continuity properties of the considered policies.
268 11 Controlled Stochastic Differential Equations

In deterministic optimal control theory, relaxed controls [322, 340] allow for the mathematical analysis on continuity-
compactness to be applied with no artificial restrictions on the classes of control policies considered. A particularly conse-
quential approach is via the study of topologies on Young measures defined by randomized/relaxed controls [227,340], [74,
Section 2.1], [322, p. 254], [223] where one views the topology on control policies to be identified with the weak conver-
gence topology of a measure defined on a product space with a fixed marginal at an input/state space (typically the Lebesgue
measure in optimal deterministic control).
Let us consider an open-loop controller, where the control is only a function of the time variable. We let ν(dt, du) be
a measure on [0, T ] × U where the first marginal λ(dt) is the normalized Lebesge measure on time interval [0, T ] and
let ν(du|t) = 1{γ(t)∈du} be the conditional measure induced by deterministic open loop control. So, any deterministic
open-loop control is embedded via:
ν(dt, du) = λ(dt) 1{γ(t)∈du} .
If allow for randomized policies, we obtain the set Pλ ([0, T ] × U) of all probability measures with fixed marginal on
[0, T ]. This set is weakly closed, whose extreme points are those induced by deterministic policies. Thus, any deterministic
optimal control problem, which can be written in an integral form and have lower semi-continuous cost functions in actions,
will have an optimal solution, which will then be deterministic as these form the extreme points of randomized controls.
It can also in fact be shown that such policies are dense in the space of randomized policies, in addition to these policies
forming the extreme points in the set of randomized policies (see e.g., [32, Proposition 2.2] [209], [237, 19, Theorem 3],
but also many texts in optimal stochastic control where denseness of deterministic controls have been established inside
the set of relaxed controls [54]). We refer the reader to [223] for further discussion.
The following example builds on these, with somewhat different arguments. Let X = R, U = [0, 1], and let f : X × U →
[0, 1] and c : X × U → [0, 1] be measurable functions continuous in the control action variable. Consider the following
optimal control problem:
Z 1
inf c(xt , ut ) λ(dt) (11.31)
γ:X→U 0
ut =γ(xt )

subject to

dx
= f (xt , ut ) (11.32)
dt
The natural space to consider is the set of all control functions which depends on the current state, where the only restriction
is measurability. However, allowing for measurability only does not facilitate continuity/compactness arguments since,
as noted above, imposing compactness on a space of functions is an unnecessarily restrictive condition. Accordingly,
one often cites appropriate but tedious measurable selection theorems building on optimality equations through dynamic
programming.
On the other hand, every deterministic function of state can be expressed as a deterministic function of time, and so, be
considered open-loop. Accordingly, we consider open loop controls and those which are relaxed. Let Pλ ([0, T ] × U) be the
set of relaxed open loop policies (known as Young measures). Now consider the space C([0, 1]; X)×Pλ ([0, T ]×U), where
C([0, 1]; X) is the space of continuous functions from [0, 1] to X. We endow this space with the product topology with the
first component being under the supremum norm and the second under Prohorov metric (or any weak convergence inducing
metric). Note now that the cost (11.35) is continuous on C([0, 1]; X) × Pλ ([0, T ] × U). Note that since f is uniformly
bounded, we have that the set S of all admissible sample paths of the state x : [0, 1] → X is equicontinuous, and so, by the
Arzela-Ascoli theorem, S is relatively compact in C([0, 1]; X). Accordingly, our space of interest S × Pλ ([0, T ] × U) is a
relatively compact subset of C([0, 1]; X) × Pλ ([0, T ] × U).
Define now
 Z t 
H= (x, m) ∈ C([0, 1]; X) × Pλ ([0, T ] × U) : xt − f (xs , u) ms (du) λ(ds) = 0 , (11.33)
0

where ms (du) = m(du|s). This set is closed under the topology defined on C([0, 1]; X) × Pλ ([0, T ] × U) and is a subset
of C([0, 1]; X) × Pλ ([0, T ] × U). Hence, H is compact. Now, the problem then is to find an optimal (x, m) ∈ H which
11.5 Existence of Optimal Policies under Full, Partial and Decentralized Information 269

minimizes (11.35), reformulated as: Z 1


inf c(xt , u) m(dt, du)
(x,m)∈H 0

This is continuous in (x, m) by an application of the generalized weak convergence theorem under continuous convergence
[287, Theorem 3.5] or [211, Theorem 3.5]. Therefore, there exists an optimal solution to the problem.

11.5.2 Existence of Optimal Policies for Fully Observed Stochastic Models

We start the discussion with the fully observed model. We build on [202], [204]) and [258]
Consider a continuous-time process Xt taking values in a Euclidean space RN , controlled by a control process {Ut } taking
values in a compact metric space U. In the context of a diffusion process,

dXt = b(Xt , Ut )dt + σ(Xt )dWt , (11.34)

driven by standard Brownian motion {Wt }, under the control policy U and the initial condition x ∈ Rd . We also allow the
control policy to be randomized, that is P(U)-valued, where P(U) denotes the space of probability measures on U under
the weak convergence topology. An admissible control is a P(U) valued non-anticipative process {Ut } .
To ensure existence and uniqueness of weak solutions of (11.34), we impose the following assumptions on the drift b and
the diffusion matrix σ .
(A1)The function b is jointly continuous in (x, u) and σ is locally Lipschitz continuous, i.e., for some constant CR > 0
depending on R > 0, we have
2
∥σ(x1 ) − σ(x2 )∥ ≤ CR |x1 − x2 |2
p
for all x1 , x2 ∈ BR , where ∥σ∥ := (σσ T ) . Also, we assume that b, σ are uniformly bounded, i.e.,

sup |b(x, u)| + ∥σ(x)∥ ≤ C ∀ x ∈ RN ,


u∈U

for some constant C > 0.


(A2)For some Ĉ1 > 0, it holds that
d
X
aij (x)zi zj ≥ Ĉ1 |z|2 ∀ x ∈ RN ,
i,j=1

and for all z = (z1 , . . . , zd ) ∈ RN , where a := 21 σσ.


In view of (A2), one sees that σ −1 exists and it is bounded . For similar existence/approximation results in the fully
observable setup, the authors in [202, 204, 207], assumed that b, σ are bounded and uniformly Lipschitz.
In the following, first we will trace some of the ideas presented by Kushner (see e.g. [207]), though with some presentational
differences and then present an alternative approach. Suppose that one wishes to minimize the cost
Z T 
U
J(U ) := Ex c(Xs , Us )ds + cT (XT ) , (11.35)
0

over all admissible control policies. Here, c, cT are continuous and bounded functions. We define a relaxed wide-sense
admissible control policy in the following. We first place the Young topology on the control action space, by viewing the
progressively measurable random control process m(dt, du)(ω) to be a random probability measure on [0, T ] × U with
its fixed marginal on [0, T ] to be the Lebesgue measure; here Pλ ([0, T ] × U) (the space of such probability measures) is
endowed with the weak convergence topology. We require that m[0,t] be independent of Bs −Bt , s > t for every t ∈ [0, T ].
We let m ∈ P(Pλ ([0, T ] × U)). We then consider the C([0, T ]; RN )-valued (under sup-norm) Xt process (solution to the
diffusion equation (11.34)) induced by m(dt, du)(ω) and then consider the space of probability measures on these random
variables.
270 11 Controlled Stochastic Differential Equations

From [15, Theorem 2.2.11], it is easy to see that under any choice of control process m(dt, du)(ω), (11.34) admits a unique
weak solution .
Toward an existence and approximation analysis, we adopt the following two approaches.

Weak convergence approach without measure transformation

In one approach, presented extensively by several seminal studies by Kushner and collaborators [202, 204, 207] as well as
others such as Borkar [55], one considers the following. Given the above, we consider

H = (η, m) ∈ P(C([0, T ]; RN )) × P(Pλ ([0, T ] × U)) :
 Z t 
Ex f (Xt ) − f (X0 ) − Au f (Xs )ms (du)λ(ds) = 0,
0

m[0,t] is independent of Ws − Wt , s > t, s, t ∈ [0, T ] , (11.36)

for all twice continuously differentiable function f with compact support and where ms (du) = m(du|s) and

Au f (x) := a(x)∇2 f (x) + b(x, u) · ∇f (x) .




In the following theorem we show that the space H is closed, we follow Kushner’s weak convergence approach (see
e.g., [202, 204, 207]) but under weaker conditions. For detailed proof see [258, Theorem 2.1] .

Theorem 11.5.1 [258] Suppose that Assumptions (A1)–(A2) hold. Then, if mn → m, with (ηn , mn ) ∈ H, the measure on
the state process η n → η which is the measure on the state process under m (that is, (η, m) ∈ H). Thus, H is closed.

Now, the problem is to find an optimal (η, m) ∈ H which minimizes (11.35). From this, one can, as in the deterministic
case summarized in [271, Section 7.1], arrive at general conditions for the existence of an optimal solution. One can also
establish conditions (see e.g., [202, Theorem 4.4]) for compactness for this set under the weak topology. We can write the
cost as
 Z T Z  
J(m) = E c(Xs , u)ms (ω)(du) + cT (XT ) . (11.37)
0 U

Theorem 11.5.2 [258] Suppose that Assumptions (A1)–(A2) hold. Then,

J : P(Pλ ([0, T ] × U)) → R

is a continuous map.

Next, using the above continuity result we want to prove the near optimality of piece-wise constant policies .

Theorem 11.5.3 [258] Suppose that Assumptions (A1)–(A2) hold. Then, for every ϵ > 0, there exists a piece-wise constant
control policy in ΓRC (thus also non-anticipative) which is ϵ-optimal.

Proof. From [15, Theorem 2.3.1], we know that the set of non-anticipative measures with quantized support (in both time
and control) is dense in Pλ ([0, T ]×U) . Thus, by the continuity of the cost as a function of policies (as we have established
in Theorem 11.5.2), we obtain our result .
11.5 Existence of Optimal Policies under Full, Partial and Decentralized Information 271

An approach with measure transformation

In an alternative approach, we consider the state process to be exogenous and the control only impacting the cost function.
For the analysis of this subsection, we are assuming that the running cost c(x, u) is bounded measurable and continuous
in its second argument (i.e., only in u) and cT is bounded measurable. differently from the previous subsection, we define
relaxed wide-sense admissible control policy in the following. As in the above, we first place the Young topology on the
control action space, by viewing the control to be a probability measure on C([0, T ]; RN ) × Pλ ([0, T ] × U) with its fixed
marginal on C([0, T ]; RN ) to be the Wiener measure, moreover we require that, under the measure-transformed model,
m[0,s] be independent of Wt − Ws for any t > s .
Let ΓWRC denote the space of all wide-sense admissible control policies . A typical element of ΓWRC is denoted by m
(without loss of generality). To study this, we adopt Girsanov’s measure transformation. Define

dXt′ = σ(Xt′ )dWt ,

which, in fact, is policy independent (under new probability measure P0 ), where


Z T
1 T −1

dP
Z
=: ZT = exp σ −1 (Xs )b(Xs , Us )dWs − |σ (Xs )b(Xs , Us )|2 ds .
dP0 0 2 0

provided that b is integrable (uniform over control policies) and σ −1 (x) exists and is bounded (which is a consequence of
(A2)) . In this case, the marginal on the state process is fixed, but the cost function is now represented as:
 Z T 
U dP
J(m) := EP0 c(Xs , Us )ds + cT (XT )
dP0 0

Here, one wishes to minimize (where the measure on the path process is fixed)

inf J(m) . (11.38)


m∈ΓWRC

In the following, we adopt the latter approach, as it will be much simpler to be generalized to information structures beyond
the fully observed model, including decentralized information structures. This analysis is based on the supporting result in
Lemma 11.5.2. Which shows that the Radon-Nikodym derivative is continuous as a function of policies (under a suitable
topology over the policy space) .

Lemma 11.5.1 The space ΓWRC under the weak convergence topology is compact.

Proof. Note that ΓWRC ⊂ P C([0, T ]; RN ) × Pλ ([0, T ] × U) . Since (in ΓWRC ) the marginal on C([0, T ]; RN ) is fixed
and Pλ ([0, T ] × U) is compact (via Prohorov’s theorem), it follows that ΓWRC is tight. Thus relatively compact by Pro-
horov’s theorem. Since independence is preserved under weak convergence of probability measures (see e.g. the proof
of [131, Lemma 2.3] or [346, Theorem 5.6]), thus ΓWRC is also closed, hence compact .
The next lemma shows that the Radon-Nikodym derivative is continuous as a function of policies over ΓWRC (under the
topology of weak convergence) .

Lemma 11.5.2 Suppose that Assumptions (A1)–(A2) hold. Then, on ΓWRC , the map
Z T
U 7→ exp σ −1 (Xs )b(Xs , Us )dWs
0
Z T 
1
− |σ −1 (Xs )b(Xs , Us )|2 ds
2 0

is continuous in L1 norm.
272 11 Controlled Stochastic Differential Equations

Using this continuity property of the Radon-Nikodym derivative as a function of policy, we arrive at the following conti-
nuity result. The proof of the following lemma follows from [258, Lemma 2.6] .

Lemma 11.5.3 Suppose that Assumptions (A1)–(A2) hold. Then, J is continuous in m ∈ ΓWRC under the weak convergence
topology.

Next theorem proves the existence of an optimal policy in ΓWRC . Also, it shows that the piece-wise constant policies in
ΓWRC are near optimal .

Theorem 11.5.4 [258] Suppose that Assumptions (A1)–(A2) hold. Then, we have
(i) There exists an optimal control policy in ΓWRC .
(ii) For every ϵ > 0, there exists a piece-wise constant control policy in ΓWRC (thus also non-anticipative) which is
ϵ-optimal.

11.5.3 Existence of Optimal Policies for Partially Observed Models

Consider now a partially observed continuous-time process {Xt } on RN , controlled by a control process {Ut } taking
values in a compact action space U ⊂ RL , and with an associated observation process {Yt } taking values in RM , where
0 ≤ t ≤ T . The evolution of {Xt , Yt } is given by the stochastic differential equations

dXt = b(Xt , Ut )dt + σ(Xt )dWt ,


dYt = g(Xt )dt + dBt , (11.39)

where, g : RN → RM is a continuous and bounded function and W and B are independent standard Wiener processes
with values in RN and RM , respectively (hence, σ is a N × N -matrix). The objective is to minimize the following cost
function
Z T 
E c(Xt , Ut )dt + cT (XT ) , (11.40)
0

where c : RN × U → [0, ∞) and cT : RN → [0, ∞) are bounded and continuous functions.


The idea is again to first apply Girsanov’s transformation so that the measurements Yt form an independent Wiener process
under new probability measure Q . Following Fleming and Pardoux [131, p. 264], we define an admissible control as a
probability measure on C([0, T ]×RM )×Pλ ([0, T ]×U) with its fixed marginal on C([0, T ]×RM ) be the Wiener measure.
In addition, under the new measure Q, Yr −Yt is independent of {X0 , W· , Ys , ms ; s ≤ t}, for any 0 ≤ t ≤ r ≤ T . Let ΓWS
denote the space of such policies, where we endow this space with the weak convergence topology. As inc Lemma 11.5.1,
we have ΓWS is compact under weak convergence topology . Without loss of generality, a typical element of ΓWS is denoted
by m.
As in Section 11.4, suppose that g ≡ 0, that is Yt is just an independent process. Let us call this probability measure on the
processes as Q := PX × QY .
Now, consider again our original process where g is non-zero. Let P be the joint probability measure on the state and the
measurement processes. Since P ≪ Q, we have that
Z
f (X[0,t] , Y[0,t] )P(dX[0,t] , dY[0,t] )
Z
= G(X[0,t] , Y[0,t] )f (X[0,t] , Y[0,t] )Q(dX[0,t] , dY[0,t] )

for some Q-integrable function G, which is the Radon-Nikodym derivative of P with respect to Q. Under mild conditions,
we have that
11.5 Existence of Optimal Policies under Full, Partial and Decentralized Information 273

dP
Gt := G(X[0,t] , Y[0,t] ) = (X[0,t] , Y[0,t] )
dQ
Rt Rt
g(Xs )dYs − 21 |g(Xs )|2 ds
=e 0 0 ,
Rt
with 0 g(Xs )dYs being a stochastic integral, this time with respect to the random measure/process Ys . This relation
allows us to view the partially observed problem as one with independent measurements, with the dependence pushed to
the Radon-Nikodym derivative. In particular, we get an equivalent model

dXt = b(Xt , Ut )dt + σ(Xt )dWt ,


dYt = dBt . (11.41)

Theorem 11.5.5 [258] Suppose that the drift term b and the diffusion matrix σ satisfy Assumptions (A1)–(A2), uniformly
with respect to y ∈ RM . Then,
(i)
 T 
dP
Z
J(m) = E (X[0,T ] , Y[0,T ] )( c(Xt , Ut )dt + cT (XT )) , (11.42)
dQ 0

is continuous over the space of wide-sense admissible policies ΓW S .


(ii) There exists an optimal control policy in ΓW S .
(iii)For every ϵ > 0, there exists a piece-wise constant control policy in ΓW S which is ϵ-optimal.

Accordingly, we can again establish both existence and discrete approximation results. Once again, the above will allow
us to approximate a continuous-time process with a (sampled) discrete-time process and the machinery developed for
discrete-time optimal control will be applicable.

Remark 11.16. The utility of this approach was already observed in Section 10.8 (see Remark 10.17). In particular, if
one makes the measurements independent, so that the information structure is first static, and then makes the informa-
tion structure classical by considering the actions at time t measurable on the filtration generated by the past noise pro-
cesses and actions up to time t; Theorem 10.16, building on [346, Theorem 5.6], can be adapted to show that such a
set of measurement-action measures (with fixed marginal on the measurements) that satisfy conditional independence
u[0,t] ↔ y[0,t] ↔ ys − yt ; (x0 , W[ 0, T ]) is weakly closed. Furthermore, the value is continuous in the joint measure on
{xs ; (u, y)s , s ∈ [0, T ]} and this set of measures is tight. These lead to the compactness-continuity conditions and accord-
ingly an existence result for optimal policies follows. Furthermore, by showing that the set of {(u, y)s , s ≥ 0} measures
which have quantized support in the measurement variable are dense, one can show also that piece-wise constant control
policies are nearly optimal. This allows one to approximate a continuous-time process with a (sampled) discrete-time pro-
cess and the machinery developed earlier in the lecture notes are applicable. This approach is the essence of Kushner’s
method [204, p. 278] [202], though stated somewhat differently. This approximation result by discrete-time models also
applies for fully-observed models with a similar argument (see Exercise 11.10.2(b)).

Remark 11.17. It may be important to note that Bismut [46] arrived at further existence results, through an approach which
avoids separation (and the construction of a belief-MDP), in discrete-time a similar approach is given in Section 10.8.1.

Remark 11.18 (Revisiting the Discrete-time Case). Inspired by the work of Fleming and Pardoux [131], Borkar introduced
wide-sense control policies to study discrete-time partially-observed finite state-observation Markov decision processes
with average cost criterion (see [58, 60, 62, 63]). We recognize also that Borkar achieves what is in essence equivalent to
Witsenhausen’s static reduction reviewed earlier in Section 10.4.2. For simplicity, we only consider here the case where
state and observation spaces are finite. We consider a discrete-time Markov decision process {xn } on a finite state space X,
controlled by a control process {un } taking values in a compact Borel action space U, and with an associated observation
process {yn } taking values in a finite observation space Y, where n = 0, 1, 2, . . .. The evolution of {xn , yn } is given by
274 11 Controlled Stochastic Differential Equations

P xn+1 , yn+1 ∈ · xm , xm , um , m ≤ n = ρ(xn+1 , yn+1 ∈ ·|xn , un ),

where ρ : X × U → P(X) × P(Y) is some transition kernel. To ease the exposition, we assume that ρ is of the following
form:

ρ(xn+1 , yn+1 |xn , un ) = r(yn+1 |xn+1 ) ⊗ p(xn+1 |xn , un ),

where p is the state transition kernel and r is the observation kernel. The initial distribution of x0 is µ.
A control process {un } is admissible in classical sense if it is adapted to the filtration {σ(ym , m ≤ n)} generated by
observations {yn }. In this case, one can write

un = πn (y0 , . . . , yn ), n ≥ 0, (11.43)
Qn
for some πn : k=0 Y → U. Let us denote π = {πn }.
Note that one can always write the evolution of the state process {xn } as a noise-driven dynamical system

xn+1 = F (xn , un , wn ), (11.44)

where F : X × U × [0, 1] → X is measurable and {wn } are independently and identically distributed uniformly on [0, 1].
Using this dynamical system, we now reproduce the above process on a more convenient probability space. This will then
enable us to define wide-sense admissible policies.
We can thus reduce the problem to an independent static one via Witsenhausen/Girsanov/Borkar, see Borkar’s [58, 62]
explicit analysis or Witsenhausen’s method presented in Section 10.4.2.
Under this reduction, we obtain a new probability space P0π under which:
(a) {yn } is i.i.d. uniform on Y and independent of x0 and {wn },
(b) {u0 , . . . , un , y0 , . . . , yn } is independent of {wn }, x0 , and {ym , m > n}, for all n.
Using these properties, Borkar defines P0 to be wide sense admissible if P0 satisfies (a) and (b). Such a notion allows for
closedness of conditional independence properties under weak convergence of joint probability measures, and thus leads
to very general existence results. See [271] for a subtle clarification.

11.5.4 Existence for Models with Decentralized Information

Decentralized Model with Local Measurements

Consider now a continuous-time process {Xt } on a Euclidean space Rn , controlled by a collection of control process
Ut := {Utk , k = 1, · · · , N } with each Utk taking values in a compact Borel action space Uk ⊂ RL , and with an asso-
ciated observation process {Ytk } taking values in RM , where 0 ≤ t ≤ T . Let Y = {Y 1 , · · · , Y N }. The evolution of
{Xt , Ytk , k = 1, · · · , N } is given by the stochastic differential equations

dXt = b(Xt , Ut1 , · · · , UtN )dt + σ(Xt )dWt ,


dYti = g i (Xt )dt + dBti , i = 1, · · · , N. (11.45)

Where, W and B i , i = 1, · · · , N are independent standard Wiener processes with values in Rn and RM , respectively and
g i : Rn → Rn is a continuous and bounded function . In this section, we assume that the drift term b and the diffusion
matrix σ satisfies similar conditions as in (A1) and (A2). In particular, they satisfy the following:
QN  
(D1)The function b : Rn × (RM )N × k=1 Uk → Rn is jointly continuous and σ = σ ij : Rn × (RM )N → Rn×n
is locally Lipschitz continuous in x (uniformly with respect to the y ∈ (RM )N ). In particular, for some constant
CR > 0 depending on R > 0, we have
11.5 Existence of Optimal Policies under Full, Partial and Decentralized Information 275
2
∥σ(x1 , y) − σ(x2 , y)∥ ≤ CR |x1 − x2 |2
p
for all x1 , x2 ∈ BR , y ∈ (RM )N , where ∥σ∥ := (σσ T ) .

(D2)The functions b and σ are uniformly bounded, i.e., for some constant C > 0,
2
sup |b(x, y, u)| + ∥σ(x, y)∥ ≤ C ∀ x ∈ Rd , y ∈ (RM )N .
QN
u∈ Uk
k=1

(D3)For some Ĉ1 > 0, it holds that


d
X
aij (x, y)zi zj ≥ Ĉ1 |z|2 ∀ x ∈ Rn , y ∈ (RM )N ,
i,j=1

and for all z = (z1 , . . . , zd ) ∈ Rn , where a := 21 σσ T .


The objective here is to minimize the following cost function
Z T 
1 N
E c(Xt , Ut , · · · , Ut )dt + cT (XT ) , (11.46)
0
QN
where c : Rn × k=1 Uk → [0, ∞) and cT : Rn → [0, ∞) are bounded continuous functions. For similar existence
analysis the authors in [75, 79–81] assumed that the b, σ are uniformly Lipschitz continuous .
We define the following decoupled measurement model.

dXt = b(Xt , Yt , Ut1 , · · · , U N )dt + σ(Xt , Yt )dWt ,


dYti = dBti , i = 1, · · · , N. (11.47)

Let Y = {Y 1 , · · · , Y N } and U = {U 1 , · · · , U N }
QN
We let this decoupled measurement model have measure Q := PX × k=1 Qk . Let P be the joint probability measure on
the state and the measurement processes under a given policy. Since P ≪ Q, we have that
Z
f (X[0,t] , Y[0,t] )P(dX[0,t] , dY[0,t] )
Z
= G(X[0,t] , Y[0,t] )f (X[0,t] , Y[0,t] )Q(dX[0,t] , dY[0,t] ) ,

for some Q-integrable function G, which is the Radon-Nikodym derivative of P with respect to Q. Under mild conditions,
we have that
N Rt
dP
Rt i
g i (Xs )dYsi − 12 |g (Xs )|2 ds
Y
Gt := (X[0,t] , Y[0,t] ) = e 0 0
dQ i=1
Rt i
with 0 g (Xs )dYs being a stochastic integral, with respect to the random measure/process Ysi . This relation allows us to
i

view the decentralized stochastic control problem as one with independent measurements, with the dependence pushed to
the Radon-Nikodym derivative.
We define an admissible control as a probability measure on C([0, T ] × RM ) × Pλ ([0, T ] × Ui ) with its fixed marginal on
C([0, T ]×RM ) be the Wiener measure. In addition, under the new measure, Yri −Yti is independent of {X0 , W· , Ysi , mis ; s ≤
t}, for any 0 ≤ t ≤ r ≤ T and independent from all (mk· , Y·k ), k ̸= i . Let ΓDWS denote the space of such decentral-
ized wide sense admissible policies, where we endow this space with the weak convergence topology . Without loss of
generality, a typical element of ΓDWS is denoted by m = (m1 , · · · , mN ) .
276 11 Controlled Stochastic Differential Equations

Theorem 11.5.6 [258] Let Assumptions (D1)– (D3) hold. Then,


(i) Over ΓDWS (wide-sense admissible policies) the function
N
Y Rt i Rt i Z T 
g (Xs )dYsi − 12 |g (Xs )|2 ds)
J(m) =E (em 0 0 c(Xt , Ut1 , · · · , UtN )dt + cT (XT ) , (11.48)
i=1 0

is continuous.
(ii) There exists an optimal control policy in ΓDWS .
(iii)For every ϵ > 0, there exists a piece-wise constant control policy in ΓDWS which is ϵ-optimal.

Decentralized Model with Coupled Dynamics and Local State

Instead of (11.45) we now consider a collection of N agents with coupled dynamics given as

dXti = bi (Xti , Uti )dt + bi0 (Xt , Ut )dt + σ i (Xti )dBti , (11.49)
QN
for i = 1, · · · , N . Here, bi : Rn × Ui → Rn , bi0 : (Rn )N × k=1 Uk → Rn and σ i : Rn → Rn×n are given
functions and B i are independent standard Wiener processes with values in Rn , i = 1, · · · , N . We assume that bi , bi0 , σ i
for i = 1, · · · , N , satisfies the following:
ˆ For i = 1, · · · , N , we have bi : Rn × Ui → Rn and bi : (Rn )N × QN Uk → Rn are jointly continuous and for
(D1) 0 k=1
some constant CR > 0 depending on R > 0, we have
2
σ i (x1 ) − σ i (x2 ) ≤ CR |x1 − x2 |2

for all x1 , x2 ∈ BR .

ˆ The functions bi , bi and σ i , i = 1, · · · , N , are uniformly bounded, i.e., for some constant C > 0
(D2) 0

2
sup |bi (x, u)| + σ i (x) ≤ C ∀ x ∈ Rn ,
u∈Ui

and supu∈QN Uk
|bi0 (x, u)| ≤ C for all x ∈ (Rn )N .
k=1

ˆ For some Ĉ1 > 0, it holds that


(D3)
d
X
ak,ij (x)zi zj ≥ Ĉ1 |z|2 ∀ x ∈ Rn , k = 1, · · · , N ,
i,j=1

and for all z = (z1 , . . . , zd ) ∈ Rn , where ak := 21 σ k (σ k ) .


ˆ it is easy to see that (σ i )−1 exists and is bounded for all i = 1, · · · , N .
In view of (D3),
The objective here is to minimize the following cost function
Z T 
E c(Xt , Ut1 , · · · , U N )dt + cT (XT ) , (11.50)
0
QN
where c : (Rn )N × k=1 Uk → [0, ∞) and cT : (Rn )N → [0, ∞) are bounded continuous functions . We assume that the
control policies are only locally measurable, that is Uti is measurable with respect to σ(X[0,t]
i
) for all t ∈ [0, T ].
We define the following decoupled (non-interacting) agent model.
11.6 Near Optimality of Control Policies Designed for Discrete-time Models via Sampling 277

dXti = σ i (Xti )dWti , i = 1, · · · , N ; (11.51)



ˆ Rt
let us have the driving noise process Wt1 + 0 (σ i )−1 (Xs1 ) b1 (Xs1 , Us1 ) + b10 (Xs , Us ) ds, · · · , W

Since σ i is invertible (follows from (D3)),

R t N −1 N
have measure µ and the independent process (Wt1 , · · · , WtN ) have
N N N N

0
(σ ) (X s ) b (X s , U s ) + b 0 (X s , U s ) ds

measure µ0 . Let b̂i (Xs , Us ) = b1 (Xs1 , Us1 ) + b10 (Xs , Us ) . Then, by Girsanov, we know that the density for µ0 with


respect to µ is
N
dµ0
R T i −1 R T i −1
( (σ ) (Xt )b̂i (Xt ,Ut )dWsi − 12 (σ ) (Xt )b̂i (Xt ,Ut )2 ds)
Y
= e 0 0 .

k=1

As in Section 11.5.2, we define a relaxed wide-sense admissible control policy by first placing the Young topology on the
control action space, by viewing the control process to be a probability measure on C([0, T ]; Rd ) × Pλ ([0, T ] × Ui ) with its
fixed marginal on C([0, T ]; Rd ) to be the Wiener measure . We require also that mi[0,t] be independent of Wsi − Wti , s > t
for every t ∈ [0, T ] and independent from mj , Wsj , j ̸= i, s ∈ [0, T ]. We call again such policies decentralized locally
wide sense admissible policies, and denote with ΓDWS . Without loss of generality, a typical element of ΓDWS is denoted by
m = (m1 , · · · , mN ) . Also, it is easy to see that in the trasformed model, the measure on the path space is fixed .
Now, following the analysis in Theorem 11.5.4, 11.5.6, we have the following theorem.

ˆ (D3)
Theorem 11.5.7 [258] Suppose that Assumptions (D1)– ˆ hold. Then,

(i) Over the space of wide-sense admissible policies ΓDWS the function
 Z T 
J(m) = E m c(Xt , Ut1 , · · · , U N )dt + cT (XT )
0
N RT R T i −1 
(σ i )−1 (Xt )b̂i (Xt ,Ut )dWsi − 12 (σ ) (Xt )b̂i (Xt ,Ut )2 ds
Y
e 0 0 (11.52)
i=1

is continuous.
(ii) There exists an optimal control policy in ΓDWS .
(iii)For every ϵ > 0, there exists a piece-wise constant control policy in ΓDWS which is ϵ-optimal.

11.6 Near Optimality of Control Policies Designed for Discrete-time Models via Sampling

In view of the results presented in the previous section, we have that piece-wise constant policies are near optimal. These
then lead to discrete-time models whose solutions will be near optimal, and applicable to the original problems. For each
of the information structures below, we will consider the following arguments: (i) We obtain a sequence of discrete-time
models Tn (here the subscript n denotes the index of the model sequence, and not the fixed dimension of the state space)
arrived from piece-wise constant control policies applied to the true model T . (ii) We show that the solution to the optimal
discrete-time model J(Tn ) leads to a solution which is near optimal. The direction, limn→∞ J(Tn ) ≤ J(T ) + ϵ follows
from the analysis above and the direction limn→∞ J(Tn ) ≥ J(T ) follows from the fact that restricting to piece-wise
constant policies cannot lead to a better policy when compared with arbitrary admissible policies. (iii) We then show that
the policy obtained to solve J(Tn ) can be applied to T (and thus the approach is constructive) and is near optimal for large
n.
278 11 Controlled Stochastic Differential Equations

11.6.1 Fully Observed Setup

Consider the fully observed setup discussed in Section 11.5.2 with model (11.34) and cost criterion (11.35). That is, with
dynamics
dXt = b(Xt , Ut ) + σ(Xt )dBt ,
RT
and cost criterion J(U ) = E[ 0 c(Xs , Us )ds + cT (XT )].
Discrete-Time Model to be Solved for Near Optimal Solutions. Let Xh , Uh be the solution of the sampled process
corresponding to (11.34) with piece-wise constant policies: Uh,s = Ukh for kh ≤ s < (k + 1)h . We have that
Z (k+1)h Z (k+1)h
X(k+1)h = Xkh + b(Xs , Ukh )ds + σ(Xs )dWs (11.53)
kh kh

T
The cost can be written as, with Nh = h,

h −1
NX
E[ ĉ(Xkh , Ukh ) + cT (XNh h )] (11.54)
k=0

Rh
where ĉ(x, u) = E[ 0 c(Xs , U0 )ds|X0 = x, U0 = u] . With n = h1 , the above define a discrete-time MDP with transition
kernel Tn , cost function ĉn and total cost Jn (U ).
Thus, one defines a discrete-time model in which Xk := Xkh and Uk := Ukh for k ∈ Z+ .
The information structure at time t contains the continuous-time measurements. However, since for such fully observed
model, Markov policies are optimal, it suffices to the controller to only use the discrete-time measurements.

Theorem 11.6.1 [258] Suppose that Assumptions (A1)–(A2) hold. Then, the value of the discrete-time model convergences
to the value of the original continuous-time model. Moreover, for every ϵ > 0, there exists h > 0 so that the solution of the
discrete-time approximation gives a policy which is near optimal for the original continuous-time model.

The above apply also to the partially observed and decentralized models.

Partially Observed Setup

Consider the setup in Section 11.5.3 with dynamics (11.39) and criterion 11.40.
Discrete-Time Model to be Solved for Near Optimal Solutions. Let us have a piece-wise constant control policy with
Uh,s = Ukh for kh ≤ s < (k + 1)h . Let Xh , Yh , Uh be the solution of the sampled process corresponding to (11.34) with
piece-wise constant policies. We have that
Z (k+1)h Z (k+1)h
X(k+1)h = Xkh + b(Xs , Ukh )ds + σ(Xs )dBs
kh kh
Z t Z t
Yt = Ykh + g(Xs )ds + dBs , t ∈ [kh, (k + 1)h) (11.55)
kh kh

Thus, one defines a discrete-time model in which Xk := Xkh and Uk := Ukh for k ∈ Z+ , and the path-valued discrete-time
measurement Ȳk = Y[kh,(k+1)h) , for k ∈ Z+ .
T
The cost can be written as, with Nh = h,

h −1
NX
E[ ĉ(Xkh , Ukh ) + cT (XNh h )] (11.56)
k=0
11.6 Near Optimality of Control Policies Designed for Discrete-time Models via Sampling 279
Rh
with ĉ(x, u) = E[ 0 c(Xs , U0 )ds|X0 = x, U0 = u] . With n = h1 , the above define a discrete-time POMDP with transition
kernel Tn and cost function ĉn . Following the proof technique as in Theorem 11.6.1 (and Theorem 11.5.4), we obtain the
following near- optimality result for partially observable model .

Theorem 11.6.2 [258] Suppose that the drift term b and the diffusion matrix σ satisfy Assumptions (A1)–(A2), uniformly
with respect to y ∈ RM . Then, the optimal value of the discrete-time model convergences to the optimal value of the
original continuous-time model. Moreover, for every ϵ > 0, there exists h > 0 so that the solution of the discrete-time
approximation gives a policy which is near optimal for the original continuous-time model.

The result above, however, while involves discrete-time control and state, requires having access to the path-valued mea-
surements. It would be desirable to obtain a discrete-time model with discrete-time measurements Ykh in the original
measurement space. That is, at time kh, we would like to have Ukh = γ(Yih , i ∈ {0, 1, · · · , k}) under an admissible
policy γ. The following result (refinement to Theorem 11.6.2), building on Lusin’s theorem, achieves this (for more details
see [258, Section 5.2]).

Theorem 11.6.3 [258] Suppose that the drift term b and the diffusion matrix σ satisfy Assumptions (A1)–(A2), uniformly
with respect to y ∈ RM . Then, the optimal value of the discrete-time model (11.55) convergences to the optimal value of
the original continuous-time model. Moreover, for every ϵ > 0, there exists h > 0 so that the solution of the discrete-time
ϵ
approximation gives a policy (i.e., a policy Ukh = γ ϵ (Yih , i ∈ {0, 1, · · · , k}) obtained in the deiscretized molde of (11.41)
as in (11.58)) which is near optimal for the original continuous-time model.

Decentralized Setup

Decentralized Model with Local Measurements. Consider Section 11.5.4 with dynamics (11.45) and cost criterion
(11.46). In particular the model is

dXt = b(Xt , Ut1 , · · · , U N )dt + σ(Xt )dWt ,


dYti = g i (Xt )dt + dBti , i = 1, · · · , N. (11.57)

Discrete-Time Model to be Solved for Near Optimal Solutions. Let us have a piece-wise constant control policy with
i i
Uh,s = Ukh for kh ≤ s < (k + 1)h, i = 1, · · · , N . Let Xh , Yhi , Uhi be the solution of the sampled process corresponding
to (11.57) with piece-wise constant policies. We have that
Z (k+1)h Z (k+1)h
1 N
X(k+1)h =Xkh + b(Xs , Ukh ,··· , Ukh )ds + σ(Xs )dBs
kh kh
Z t Z t
i i i
Yt =Ykh + g (Xs )ds + dBsi , t ∈ [kh, (k + 1)h). (11.58)
kh kh

By the similar argument as in the partially observable case, applying Girsanov’s change of measure argument to the dis-
cretized model (11.58), we can define a new probability measure space in which the measurements of the each individual
are independent of the state process. In particular, we obtain the following equivalent discretized model
Z (k+1)h Z (k+1)h
1 N
X(k+1)h =Xkh + b(Xs , Ukh ,··· , Ukh )ds + σ(Xs )dWs
kh kh
Z t
Yti =Ykh
i
+ dBsi , t ∈ [kh, (k + 1)h) , (11.59)
kh

i
with policy Ukh = γ i (Ysi , s ≤ kh), by Lusin’s theorem, for some continuous function γci we have γ i = γci on a set of
i
measure (1−ϵi ) . Moreover, we have that the process Y[0,kh] can be approximated by its piece-wise constant interpolations.
Since both running/terminal costs are continuous, the cost (11.60) under a policy γ = (γ 1 , . . . , γ N ) (under continuous-time
measurements) and its continuous approximation γc = (γc1 , . . . , γcN ) (with discrete-time measurements) are close to each
other . This enables us to obtain a discrete-time model with discrete-time measurements .
280 11 Controlled Stochastic Differential Equations

Thus, one defines a discrete-time model in which Xk := Xkh and Uki := Ukh
i
for k ∈ Z+ , and the path-valued discrete-time
i i
measurement Ȳk = Y[kh,(k+1)h) , for k ∈ Z+ .
T
The cost can be written as, with Nh = h,

h −1
NX
1 N
E[ ĉ(Xkh , Ukh , · · · , Ukh ) + cT (XNh h )] (11.60)
k=0

Rh
with ĉ(x, u) = E[ 0 c(Xs , U01 , · · · , U0N )ds|X0 = x, U0 = u] . With n = h1 , the above define a discrete-time decentral-
ized POMDP with transition kernel Tn and cost function c̃n . Again, for the decentralized model, similar proof technique
as in Theorem 11.6.1, gives us the following near- optimality result .

Theorem 11.6.4 [258] Suppose that Assumptions (D1)–(D3) hold. Then, the optimal value of the discrete-time model
(11.58) convergences to the optimal value of the original continuous-time model. Moreover, for every ϵ > 0, there exists h >
i,ϵ
0 so that the solution of the discrete-time approximation gives a policy (i.e., a policy Ukh = γ i,ϵ (Yrh , r ∈ {0, 1, · · · , k}),
i = 1, · · · , N an optimal solution of (11.59)) which is near optimal for the original continuous-time model.

Coupled Dynamics and Local State. An analogous result is applicable for the model in Section 11.5.4.

11.6.2 An alternative discrete-time approximation: Euler-Maruyama discretization

In addition to the discrete-time model presented above, one can also consider the following (more standard and elementary)
Euler–Maruyama (EM) approximation for SDEs. For the given controlled SDE

dXt = b(Xt , Ut ) dt + σ(Xt ) dWt , X0 = x,

where Ut is an admissible control and Wt is a standard Brownian motion. Fix a time step h > 0 and set tk = kh. The
(explicit) Euler–Maruyama scheme under a piecewise-constant-in-time control Uth = Utk for t ∈ [tk , tk+1 ) is given by the
recursion
X̄thk+1 = X̄thk + b X̄thk , Utk h + σ X̄thk ∆Wk ,
 
(11.61)
where ∆Wk := Wtk+1 − Wtk and we denote by X̄ h (·) the usual continuous-time interpolation (piecewise constant or
piecewise linear as required).
The discrete-time per-stage cost writes as: ch (x, ζ) := c(x, ζ) × h. The associated discounted cost of the approximating
discrete-time model under the piece-wise constant control process U h is given by

Uh
X
Jα,h (x) = Ex [ β k ch (Xkh , Ukh )|X0h = x] (11.62)
k=0

for x ∈ Rd , where β = e−αh .

Let v h∗ be optimal for the discrete-time model (11.61) under cost (11.62).
Then following the arguments as in [257, Theorem 4.3], we then have near optimality of the discrete-time optimal policy
v h∗ obtained for the model (11.61) under cost (11.62) for the continuous time model as the parameter of the discretization
approaches to zero .

Theorem 11.19. Suppose that Assumptions (A1)–(A3) hold. Then we have


h∗
lim Jαv (x) = Jα∗ (x) a.e. x ∈ Rd , (11.63)
h→0
11.7 Stochastic Stability of Diffusions 281
h∗
where Jαv (x) is the expected cost induced by v h∗ for the true diffusion under the discounted cost criterion, and Jα∗ is the
value function for the diffusion under the discounted cost criterion.

The same applies for the finite horizon criterion as well.

11.6.3 Borkar’s Control Topology and a Partial Differential Equations Approach

Another approach to arrive at near optimality of discretized time and space models is via showing that the cost is continuous
on the space of control policies (stationary or Markov, depending on the cost criterion) under an appropriate topology due
to Borkar [55], and then to show that time and space quantized policies are dense in the space of all such policies [256].

11.7 Stochastic Stability of Diffusions

Recall that an Itô diffusion is a stochastic process Xt (ω) satisfying a stochastic differential equation of the form:

dXt = b(Xt )dt + σ(Xt )dBt , t ≥ s, Xs = x

where Bt is m-dimensional Brownian motion. Often b, σ satisfy regularity conditions of the form

|b(x) − b(y)| + |σ(x) − σ(y)| ≤ D|x − y|,

for some finite D (if one wishes to impose the existence of strong solutions), though this is not a requirement for the
analysis to follow. Note that here b, σ only depend on x and not on t. Thus, the process here is time-homogenous.
Continuous-time counterparts of Foster-Lyapunov criteria considered in Chapters 3 and 4 exist and are well-developed.
We refer the reader to [205], [222], [235] [234] as well as [193]. Dynkin’s formula plays a key role in obtaining the
continuous-time counterparts of the Foster-Lyapunov criteria developed in Chapter 4.
For functions V : X → R+ that are properly defined, as in the Foster-Lyapunov criteria studied in Chapter 4, conditions of
the form
AV (x) ≤ b1x∈S
AV (x) ≤ −ϵ + b1x∈S
AV (x) ≤ −f (x) + b1x∈S ,
will lead to recurrence, positive Harris recurrence and finite expectations, respectively. However, the conditions needed on
both V and the Markov process need to be carefully addressed. For example, one needs to ensure that the processes are
non-explosive, that is, they do not become unbounded in finite time; and one needs to establish conditions for the strong
Markov property. Furthermore, they must lead to a well-defined AV (x) (see Definition 11.2.7).
In the following, we review related results from Meyn and Tweedie [234, 235]. We consider processes taking values from
a locally compact Polish space X.
Let P t (x, B) := Px (Xt ∈ B) for B ∈ B(X). Let for any Borel A,
Z ∞
ηA = 1{Xt ∈A} dt,
0

denote the occupation time. We say that the process Xt is ψ-irreducible if

ψ(B) > 0 =⇒ Ex [ηB ] > 0, x ∈ X,

and the process is Harris recurrent if

ψ(B) > 0 =⇒ Px (ηB = ∞) = 1, x ∈ X.


282 11 Controlled Stochastic Differential Equations

Definition 11.20. A probability measure π on B(X) is invariant if for every B ∈ B(X)


Z
π(B) = π(dx)P t (x, B), ∀t > 0.

A Harris recurrent chain which admits an invariant probability measure is called positive Harris recurrent.
Denote by D(A) the set of all functions V : X × R+ → R for which there exists a measurable function U : X × R+ → R
such that for each x ∈ X, t > 0, Z
Ex [V (Xt , t)] = V (x, 0) + Ex [ U (xs , s)ds]

and
Z t
Ex [|U (xs , s)|]ds < ∞ (11.64)
0

In this case, we have AV (x) = U (x) and we call A the extended infinitesimal generator of the process Xt and we say that
V is in the domain of A.
In general, it is not easy to know when a function is in the domain of A. One method to enhance the set of functions that are
relevant is to consider truncated processes. Let {Om , m ∈ N} be a sequence of open bounded sets (with compact closure)
for which for every m ∈ N, Om ⊂ Om+1 ⊂ Om+2 with ∪∞ m=1 Om = X. Define:

T m = τOm
c := inf{t ≥ 0 : Xt ∈ O
c
m}

and let
ζ = lim T m .
m→∞

We call Xt non-explosive if Px (ζ = ∞) = 1 for all x ∈ X.


c
Let for m ∈ Z+ , ∆m denote a fixed state in Om and define with xm :

xm
t = xt 1{t<T m } + ∆m 1{t≥T m }

Thus, for a non-explosive process, we can define xm


t = xmin{t,T m } .
2
For Itô processes, let Am denote the extended infinitesimal generator for xm
t . In this case, Am contains C (class of
functions on X × R with continuous first and second partial derivatives).
It is important to note that in general, the domain of A may be smaller than the domain of Am , in view of the integrability
condition stated in (11.64).

Theorem 11.7.1 [235, Theorem 4.5] Let Xt be non-explosive weak Feller process: that is P t g(x) = E[g(Xt )|x0 = x] is
continuous in x for every continuous and bounded g, for all t ≥ 0. Then,
(i) If

Am V (x) ≤ −ϵ + b1x∈S , x ∈ Om , m ∈ N (11.65)

holds for some compact S, then an invariant probability exists.


(ii)

Am V (x) ≤ −f (x) + b1x∈S , x ∈ Om , m ∈ N (11.66)

holds for compact S and f : X → [1, ∞), then under any invariant probability measure π, Eπ [f (X)] ≤ b.
11.7 Stochastic Stability of Diffusions 283

Proof. Building on [132] [294, Theorem 2] (see Theorem 3.3.2 for an argument in discrete-time), for a weak Feller process
there are two possibilities: either an invariant probability exists or
Z T
1
lim sup νP s (C)ds = 0,
T →∞ ν T 0

for all compact C, where the supremum is over all initial probability measures on x0 . Condition (i) then implies that the
latter cannot take place. The second result follows from the discussion for Theorem 4.2.5 together with (i). ⋄
A useful technique in arriving at stochastic stability is to sample the process to obtain a discrete-time Markov chain, whose
stability will imply the stability of the original process through a careful construction of invariant probability measures,
similar to the discussion on sampled chains in Chapter 3: Any invariant measure for the continuous-time process is also
invariant for a sampled discrete-time process, and thus the uniqueness of an invariant measure for the sampled process
would imply the uniqueness of an invariant measure for the continuous-time process, provided one exists.
Let a be a probability measure on R+ . Define
Z
Ka (x, B) = P t (x, B)a(dt)

Thus, Ka represents a sampled chain. A Borel set S is called νa -petite if νa is a non-trivial measure and a is a probability
measure on (0, ∞) that satisfies:
Ka (x, B) ≥ νa (B), x ∈ S,
for all B ∈ B(X). Furthermore, we have the following if S is a petite set. Meyn and Tweedie define a process to be a
T -process if for some distribution a, the kernel Ka (x, A) ≥ T (x, A) where (·, A) is lower semi-continuous for each Borel
A and T (x, X) ̸= 0 for each x ∈ X. Note that strong Feller processes are T -processes as T can be taken to be Ka itself.
Recall from Theorem 3.2.8 that for an irreducible T -process, every compact set is petite [235].

Theorem 11.7.2 [235, Theorem 4.2] Let {xt } be an irreducible non-explosive process and (11.65) hold for S closed and
petite, and with V bounded on S. Then, the process is positive Harris recurrent.

We also refer the reader to [89, Theorem 4.1] and emphasize that, as in Chapter 4, irreducibility is not required for the
existence of an invariant probability measure. Theorem 11.7.2 then implies the importance of the petiteness condition on
S. As we observed earlier in Chapter 3, such sets allow for regeneration and hence lead to Harris recurrence and uniqueness
of an invariant probability measure.
Let the notation {limt→∞ |xt | = ∞} denote the event that for any compact C, for all t sufficiently large xt ∈
/ C. If

Px ( lim |xt | = ∞) = 0,
t→∞

xt is said to be non-evanescent.

Theorem 11.7.3 [235, Theorem 3.1] Let {xt } satisfy

Am V (x) ≤ b1x∈S , x ∈ Om , m ∈ N

for a compact S, b < ∞ and where V is a norm-like function (i.e., limx→∞ V (x) = ∞). Then,

Px ( lim |xt | = ∞) = 0
t→∞

for each x ∈ X.

Further stochastic stability results, beyond the existence of invariant probability measures, have found applications; for
these, we refer the reader to [205] and [312]. We state one next.
284 11 Controlled Stochastic Differential Equations

Theorem 11.7.4 [205] [312, Prop. 5.5.1] Suppose that there exists a function V : Rn → R+ which is in the domain of
Am for every m, and satisfies

AV (x) ≤ −αV (x) + b, x ∈ Om , m ∈ N (11.67)

for some α, b > 0. Then,


b
E[V (Xt )] ≤ e−αt E[V (X0 )] + ,
α
provided that E[V (X0 )] < ∞.

Exercise 11.7.1 Prove Theorem 11.7.4. Hint. Apply Dynkin’s formula to V (xt )eαt ; note A(V (xt )eαt ) = αV (xt )eαt +
eαt AV (xt ) and use (11.67).

11.8 The Wong-Zakai Theorem and Robustness of the Stratonovich Integral

The Brownian noise is an idealization and is not practically achievable. However, it is approximated arbitrarily well by
signals which are sufficiently regular, so that with these regular approximations one can define a Riemann integration.
This then raises a question of robustness and convergence of approximate integrations. The following establishes such an
approximation result.
Let {wn (t), n ∈ N} be a sequence of continuous and piece-wise differentiable in t with bounded variation approximations
which converge almost surely to a Brownian process, such that there exist random variables n0 , k so that wn (t, ω) ≤ k(ω)
almost surely and all t ∈ [a, b] when n > n0 (ω), and that wn (t) converges to Bt almost surely. In particular, we ask that
{wn (t), n ∈ N} be a sequence of continuous and piece-wise differentiable in t approximations which converges uniformly
almost surely to a Brownian process.

Theorem 11.8.1 [333] Let σ(t, x) be continuously differentiable in x, t, and let {wn (t), n ∈ N} be a sequence of approx-
imations as discussed above of a Brownian process. Then, we have that, almost surely
b b

Z Z Z
1
lim σ(t, wn (t))dwn (t) = σ(t, Bt )dt + σ(t, Bt )dBt
n→∞ a 2 a ∂t

where the first two integrations are in the Riemann sense.


Theorem 11.8.2 (Wong-Zakai Theorem) [333] Let µ(t, x), σ(t, x) be continuously differentiable in x, t and µ, σ, ∂x σ
∂ 2
be Lipschitz continuous with constant k > 0, and |σ(x, t)| ≥ β > 0 with ∂t σ(x, t) ≤ kσ (x, t). Let {wn (t), n ∈ N} be a
sequence of approximations as discussed above of a Brownian process.
If for each n ∈ N, xn is the solution to the ODE:

dxn 1 ∂ dwn
= µ(t, xn ) − σ(t, xn ) σ(t, xn ) + σ(t, xn ) , xn (a) = xa ,
dt 2 ∂x dt
almost everywhere on [a, b], then xn (t) converges almost surely uniformly in t ∈ [a, b] to a stochastic process Xt solving
the equation
dXt = µ(t, Xt )dt + σ(t, Xt )dBt , Xa = xa ,
as n → ∞.

Note that here, unlike the discrete-time approximations, we are approximating the noise as well. One way to approximate
the noise is via a piece-wise linear interpolation of discrete updates:

wn ((k + 1)h) = wn (kh) + hZk ,
11.10 Exercises 285
dwn Z
where Zk is an independent Gaussian with mean zero and variance 1. In this case, notice that dt = √
h
in between the
sampling instants.
Note also that in the above, we have the correction term − 21 σ(t, xn ) ∂x ∂
σ(t, xn ), which disappears when one considers
instead of Itô, the Stratonovich integral [174, Theorem 1.2] (see Exercise 11.10.14). With the above, we have the following.
Consider
dxn = f (xn )dwn
where the integration is in the Riemann sense. If wn (t) → B(t) as above, then the solution converges to

dx = f (x) ◦ dB,

where the integration is in the Stratonovich sense.


This last observation is yet another motivation for using the Stratonovich integral for certain applications. For related results
with control, see [203] and [255].

11.9 Bibliographic Notes

For discounted and average cost problems, analysis based on the theory of partial differential equations can be utilized
to obtain more general results [15, Chapter 3]. The regularity conditions on the value function can also be relaxed. For
stochastic integration, one can also relax conditions on the functions via Krylov’s generalization [197].
For filtering theory, the reader is referred to [203] as well as [25, 238].

11.10 Exercises

Exercise 11.10.1 a) Solve


dXt = µXt + σdBt
Hint: Multiply both sides with the integrating factor e−µt and work with d(e−µt Xt ).

b) Solve
dXt = µdt + σXt dBt
1 2
Hint: Multiply both sides with the integrating factor e−σBt + 2 σ t . Finally, verify by direct computation (via Itô’s formula)
that
σ2
Xt = X(0)eσBt +(µ− 2 )t
is the solution. The equation in b) above is often used as a model for mathematical finance where µ is called the drift and
σ is called the volatility (of the financial environment).

Exercise 11.10.2 (Girsanov’s measure transformation / static reduction) a) Consider


n
X
Sn = Xk
k=1

where Xk ∼ N (0, 1) is i.i.d. Since a sum of Gaussians is Gaussian, (S1 , · · · , Sn ) is a Gaussian random vector with
measure, say with measure Q0 , and density
 
− 21 s21 +(s2 −s1 )2 +···+(sn −sn−1 )2
Cn e
286 11 Controlled Stochastic Differential Equations
Pn
for some constant Cn . Now, instead, assume that Sn′ = k=1 (Xk + µk ) where µk is a sequence of constants. In this case,
(S1′ , · · · , Sn′ ) is a Gaussian random vector with measure Q and density, for some constant Cn′ ,
 
− 12 (s1 −µ1 )2 +(s2 −s1 −µ2 )2 +···+(sn −sn−1 −µn )2
Cn′ e
P P
1 2 2
+···+(sn −sn−1 )2 ) ( µk (Sk −Sk−1 )− 12 µ2k )
= Cn′ e− 2 (s1 +(s2 −s1 ) e k k

Thus, the Radon-Nikodym derivative is obtained as

dQ P 1
P 2
= e( k µk (Sk −Sk−1 )− 2 k µk ) (11.68)
dQ0
This is the same derivation we studied in Section 10.4.2.
With this interpretation, consider the measurement process given in (??-??) with dyt = h(xt )dt + dBt and let P be the
measure on this process. Let P0 denote the measure on an independent Brownian motion dyt′ = dBt . Now, by viewing µk
as the drift term h(xt )dt, compare (11.68) with
Z T Z T 
1
ZT = exp h(xs )dys − |h(xs )|2 ds .
0 2 0

where
dP
= ZT .
dP0

b) Repeat the above with


Sn = Sn−1 + σ(Sn )Xn ∼ Q0
where σ(·) is invertible and Sn′ = Sn−1 + µk + σ(Sn′ )Xn ∼ Q and show that in this case, we have

dQ P −1 1
P −1 2 2
= e( k σ (Sk−1 )µk (Xk )− 2 k (σ (Sk−1 )) µk ) (11.69)
dQ0

In the context of a diffusion process, dxt = h(xt )dt + σ(xt )dBt ∼ P and dx′t = σ(x′t )dBt ∼ P0 compare the above with
Z T
1 T −1

dP
Z
−1 2
=: ZT = exp σ (xs )h(xs )dBs − |σ (xs )h(xs )| ds .
dP0 0 2 0

c) In part a) if xt is a controlled process and in part b) if h takes in control as an input, the above measure transformations
then lead to stochastic analysis where one can study the problem in a new probability space where the control does not
impact the flow of the process x′t or Bt , as in static reduction, but it is dependent on them. This approach facilitates various
continuity, compactness and approximation results, leading to very general optimality results.

Exercise 11.10.3 (Feynman-Kac Formula: Expected hitting time to a boundary) Let S ⊂ Rn be a bounded open set
with smooth boundary ∂S. The following partial differential equation for the notation of the Laplacian of a function f
P ∂2f
given with ∆f := i ∂(x i )2 f (x))

1
− ∆u = 1, u ∈ S
2
u = 0 u ∈ ∂S

is known to admit a solution u(x). Now, for any initial point x ∈ S, consider the Brownian motion Bt with B0 = x. Define

τSx = inf{t ≥ 0 : Bt ∈ ∂S}


11.10 Exercises 287

Show that
u(x) = E[τSx ]

Hint. For the equation Xt = Bt , the generator satisfies the relation A(f ) = 21 ∆f . Then, with min(N, τSx ) = τ N via
Dynkin’s formula
Z τN
1
E[u(Xτ N )] = E[u(X0 )] + E[ ∆u(Xs )ds]
0 2
Since − 21 ∆u = 1 until the stopping time and u is bounded, we have that limN →∞ E[τ N ] is bounded and τSx has finite
expectation. As a result,
Z τN
1
u(x) = E[u(X0 )|X0 = x] = −Ex [ ∆u(Xs )ds] + Ex [u(Xτ N )] →N →∞ Ex [τSx ]
0 2

Finally, conclude with observing that for x ∈ ∂S u(x) = 0 (and by the above for x inside S, ∆u(x) = −1).

Exercise 11.10.4 (White Noise Property of the Brownian Noise) Let us view/define the Fourier transform of dBt to be3
defined sample path wise:
Z 1
ak (ω) = dBt (ω)e−i2πkt dt.
0
Show that ak , k ∈ Z is Gaussian, and i.i.d.

Exercise 11.10.5 Prove the Itô isometry property.

Exercise 11.10.6 Complete the details for the solution to the optimal portfolio selection problem given in Example 11.13.

Exercise 11.10.7 Solve an average-cost version of the linear quadratic regulator problem and identify conditions on the
cost function that leads to a cost that is independent of the initial condition.

Exercise 11.10.8 Consider a Brownian process in Rd . Show that this process is recurrent for d = 1, 2 but transient for
d ≥ 3.

Exercise 11.10.9 Construct an example of a control-free stochastic differential equation with additive Brownian noise
such that, while in the absence of the noise term the (deterministic) process is unstable, the presence of the noise makes the
system stochastically stable.

Exercise 11.10.10 (Finite state continuous-time Markov processes and discrete-time representation via uniformization)
Let Xt be a time-homogenous continuous-time Markov process with a finite state space X. Let for any a, b ∈ X and t ≥ 0:

P (Xt = b|X0 = x) =: Pt (x, b),

denote the continuous-time transition kernel. Define


Pt − P0
lim =: Λ
t↓0 h

to be a transition rate intensity matrix. Define P (Xt ∈ ·) = µt (·) = (µ0 Pt )(·). Then, it follows (show this as an exercise)
that

3
This integral is a well-defined Riemann-Stieltjes integral via the Young integral [339]. Note that this is not a contradiction with
Theorem 11.2.1 since what Theorem 11.2.1 implies is that the Riemann-Stieltjes integral may not be well-defined for an arbitrary
function, but if the function is sufficiently regular one can still define a Riemann-Stieltjes integration with respect to the Brownian for a
given sample path, as is the case here. See [339].
288 11 Controlled Stochastic Differential Equations

dµt Pt+h − Pt
= lim = µt Λ
dt h↓0 h
and thus
µt = µ0 eΛt
Now, define X
λ = max λi,j
i∈X
j∈X,̸=i

and define
Λ
R=I+
λ
It can be shown that R is a (discrete-time) stochastic matrix since the row-sum of Λ is zero.
Show that, if Zt is a discrete-time time-homogenous Markov chain with (discrete-time) transition kernel R, then

P (Zt = ·) = P (XNt = ·),

where Nt is an independent Poisson process with rate λ.


Hint. Show that
X Rk tk X  tk 
µt = µ0 eΛt = µ0 eλRt e−λt = e−λt = e−λt Rk
k! k!
k≥0 k≥0
 
λk tk −λt
Then, observe that P (Nt = k) = k! e and that Nt is independent from Xt , to conclude that µt = P (XNt ∈ ·).

Exercise 11.10.11 For a system with (11.9), arrive at (11.10) for the case where p is invariant. This relation is important
for the study of stochastic learning/simulated annealing problems [83].

Exercise 11.10.12 Read [217, Chapters 4 and 5] to study the maximum principle and viscosity solutions for the deterministic-
setup.

Exercise 11.10.13 Consider the diffusion process



dXt = ∇f (Xt )dt + 2dBt .

Show, via the Fokker-Planck equation, that p(x) = Kef (x) is an invariant probability measure where K is a normalizing
constant.

Exercise 11.10.14 [Itô vs. Stratonovich Equivalence] Consider

dXt = b(Xt , Ut )dt + σ(Xt ) ◦ dBt

By considering the construction of the Stratonovich integral, study that [174, Theorem 1.2]
1
σ(Xt ) ◦ dBt = σ(X(t))dBt + dσ(X(t))dBt
2
1 d 1 d
= σ(X(t))dBt + ( σ(X(t)))dX(t)dBt = σ(X(t))dBt + ( σ(X(t)))σ(Xt )dt (11.70)
2 dx 2 dx
Thus, arrive at an equivalence relation between the two integrals, so that we have

dXt = b̃(Xt , Ut )dt + σ(Xt )dBt ,

where the integral is now in the Itô sense with


11.10 Exercises 289

1 d
b̃(Xt , Ut ) = b(Xt , Ut ) + ( σ(X(t)))σ(Xt )
2 dx
12

Robustness to Incorrect Models and Learning

In many applications, typically only an ideal model (controlled transition kernel) is assumed and the control design is
based on the given model, raising the problem of performance loss due to the mismatch between the assumed model and
the actual model. In this chapter, we study continuity properties of discrete-time stochastic control problems with respect to
system models (i.e., controlled transition kernels) and robustness of optimal control policies designed for incorrect models
applied to the true system.
The chapter studies both fully observed and partially observed setups under an infinite horizon discounted expected cost
criterion as well as the average cost criterion. The results for the discounted cost criterion will also imply the identical
results for finite horizon criteria.
We will show that continuity can be established under total variation convergence of the transition kernels under mild
assumptions and with further restrictions on the dynamics and observation model under weak and setwise convergence of
the transition kernels. Using these continuity properties, we establish convergence results and error bounds due to mismatch
that occurs by the application of a control policy which is designed for an incorrectly estimated system model to a true
model, thus establishing positive and negative results on robustness. These findings entail positive implications on empirical
learning in (data-driven) stochastic control since often system models are learned through empirical training data where
typically a weak convergence criterion applies but stronger convergence criteria do not.
The chapter can also be viewed as a generalization of the approximations framework presented in Section 8.2. This con-
nection will be made explicit later in the chapter.

12.1 Introduction

We will discuss both the partially and fully observed setups. Let X ⊂ Rm denote a Borel set which is the state space of a
partially observed controlled Markov process. Let Y ⊂ Rn be a Borel set denoting the observation space of the model, and
let the state be observed through an observation channel Q. As before in the notes, the observation channel, Q, is defined as
a stochastic kernel (regular conditional probability) from X to Y, such that Q( · |x) is a probability measure on the (Borel)
σ-algebra B(Y) of Y for every x ∈ X, and Q(A| · ) : X → [0, 1] is a Borel measurable function for every A ∈ B(Y).
A decision maker (DM) is located at the output of the channel Q, and hence it only sees the observations {Yt , t ∈ Z+ }
and chooses its actions from U, the action space which is a Borel subset of some Euclidean space. As discussed earlier, an
admissible policy γ is a sequence of control functions {γt , t ∈ Z+ } such that γt is measurable with respect to the σ-algebra
generated by the information variables

It = {Y[0,t] , U[0,t−1] }, t ∈ N, I0 = {Y0 },

where
Ut = γt (It ), t ∈ Z+ , (12.1)
are the U-valued control actions and
292 12 Robustness to Incorrect Models and Learning

Y[0,t] = {Ys , 0 ≤ s ≤ t}, U[0,t−1] = {Us , 0 ≤ s ≤ t − 1}.

We define Γ to be the set of all such admissible policies. The update rules of the system are determined by (12.1) and the
following: Z

Pr (X0 , Y0 ) ∈ B = P (dx0 )Q(dy0 |x0 ), B ∈ B(X × Y),
B
where P is the (prior) distribution of the initial state X0 , and
 
Pr (Xt , Yt ) ∈ B (X, Y, U )[0,t−1] = (x, y, u)[0,t−1]
Z
= T (dxt |xt−1 , ut−1 )Q(dyt |xt ), B ∈ B(X × Y), t ∈ N,
B

where T is the transition kernel of the model. The objective of the agent (decision maker) is the minimization of the infinite
horizon discounted cost,
"∞ #
T ,γ
X
t
Jβ (c, T , γ) = EP β c(Xt , Ut )
t=0

for some discount factor β ∈ (0, 1), over the set of admissible policies γ ∈ Γ , where c : X × U → R is a Borel-measurable
stage-wise cost function and EPT ,γ denotes the expectation with initial state probability measure P and transition kernel
T under policy γ. Note that we write the infinite horizon discounted cost as a function of the transition kernels and the
stage-wise cost function since we will analyze the cost under the changes on those variables.
We define the optimal cost for the discounted infinite horizon setup as a function of the stage-wise cost function and the
transition kernels as

Jβ∗ (c, T ) = inf Jβ (c, T , γ).


γ∈Γ

Problem P1: Continuity of Jβ∗ (c, T ) under the Convergence of the Models. Let {Tn , n ∈ N} be a sequence of transition
kernels which converges in some sense to another transition kernel T and {cn , n ∈ N} be a sequence of stage-wise cost
functions corresponding to Tn which converge in some sense to another cost function c. Does that imply that

Jβ∗ (cn , Tn ) → Jβ∗ (c, T )?

Problem P2: Robustness to Incorrect Models. A problem of major practical importance is robustness of an optimal
controller to modeling errors. Suppose that an optimal policy is constructed according to a model which is incorrect: how
does the application of the control to the true model affect the system performance and does the error decrease to zero as
the models become closer to each other? In particular, suppose that γn∗ is an optimal policy designed for Tn and cn , an
incorrect model for a true model T and c. Is it the case that if Tn → T and cn → c, then

Jβ (c, T , γn∗ ) → Jβ∗ (c, T )?

Problem P3: Empirical Consistency of Learned Probabilistic Models and Data-Driven Stochastic Control. Let
T (·|x, u) be a transition kernel given previous state and action variables x ∈ X, u ∈ U, which is unknown to the de-
cision maker (DM). Suppose the DM builds a model for the transition kernels, Tn (·|x, u), for all possible x ∈ X, u ∈ U by
collecting training data (e.g. from the evolving system). Do we have that the cost calculated under Tn converges to the true
cost (i.e., do we have that the cost obtained from applying the optimal policy for the empirical model converges to the true
cost as the training length increases)?
Problem P4: Approximation by Finite MDPs as an Instance of Robustness to Incorrect Models. Can we view the
approximation problem of a continuous space MDP model with a finite model, as studied in Section 8.2, as an instance of
the robustness problem?
12.1 Introduction 293

We will study the above for the average cost criterion as well. For the average cost criterion, we have
"N −1 #
1 T ,γ X
J∞ (c, T , γ) = lim sup Ex0 c(Xt , Ut )
N →∞ N t=0

To denote the explicit dependence of the optimal cost in the transition kernel, we use the notation

J∞ (c, T ) = inf J∞ (c, T , γ).
γ∈Γ

12.1.1 Some Examples and Convergence Criteria for Transition Kernels

Convergence Criteria for Transition Kernels.

Before introducing the convergence criteria to be presented in the chapter, we refer the reader to Appendix D.

Definition 12.1.1 For a sequence of transition kernels {Tn , n ∈ N}, we say that
– Tn → T weakly if Tn (·|x, u) → T (·|x, u) weakly, for all x ∈ X and u ∈ U,
– Tn → T setwise if Tn (·|x, u) → T (·|x, u) setwise, for all x ∈ X and u ∈ U,
– Tn → T under the total variation distance if Tn (·|x, u) → T (·|x, u) under total variation for all x ∈ X and u ∈ U.

Examples [186].

Let a controlled model be given as xt+1 = F (xt , ut , wt ), where {wt } is an i.i.d. noise process. The uncertainty on the
transition kernel for such a system may arise from lack of information on F or the i.i.d. noise process wt or both:
(i) Let {Fn } denote an approximating sequence for F , so that Fn (x, u, w) → F (x, u, w) pointwise. Assume that the
probability measure of the noise is known. Then, corresponding kernels Tn converge weakly to T : If we denote the
probability measure of w with µ, for any g ∈ Cb (X) and for any (x0 , u0 ) ∈ X×U using the dominated convergence
theorem we have
Z Z
lim g(x1 )Tn (dx1 |x0 , u0 ) = lim g(Fn (x0 , u0 , w))µ(dw)
n→∞ n→∞
Z Z
= g(F (x0 , u0 , w))µ(dw) = g(x1 )T (dx1 |x0 , u0 ).

(ii) Much of the robust control literature deals with deterministic systems where the nominal model is a determinis-
tic perturbation of the actual model (see e.g. [281]). The considered model is in the following form: F̃ (xt , ut ) =
F (xt , ut )
+∆F (xt , ut ), where F represents the nominal model and ∆F is the model uncertainty satisfying some norm
bounds. For such deterministic systems, pointwise convergence of F̃ to the nominal model F , i.e. ∆F (xt , ut ) → 0,
can be viewed as weak convergence for deterministic systems by the discussion in (i). It is evident, however, that to-
tal variation convergence would be too strong for such a convergence criterion, since δF̃ (xt ,ut ) → δF (xt ,ut ) weakly
but ∥δF̃ (xt ,ut ) − δF (xt ,ut ) ∥T V = 2 for all ∆F (xt , ut ) ̸= 0 where δ denotes the Dirac measure.

(iii)Let F (xt , ut , wt ) = f (xt , ut ) + wt be such that the function f is known and wt ∼ µ is not known correctly and an
incorrect model µn is assumed. If µn → µ weakly, setwise, or in total variation, then the corresponding transition
kernels Tn converge in the same sense to T . Observe the following:
Z Z
g(x1 )Tn (dx1 |x0 , u0 ) − g(x1 )T (dx1 |x0 , u0 )
294 12 Robustness to Incorrect Models and Learning
Z Z
= g(w0 + f (x0 , u0 ))µn (dw0 ) − g(w0 + f (x0 , u0 ))µ(dw0 ). (12.2)

(a) Suppose µn → µ weakly. If g is a continuous and bounded function, then g(· + f (x0 , u0 )) is a continuous and
bounded function for all (x0 , u0 ) ∈ X × U. Thus, (12.2) goes to 0. Note that f does not need to be continuous.
(b) Suppose µn → µ setwise. If g is a measurable and bounded function, then g(· + f (x0 , u0 )) measurable and
bounded for all (x0 , u0 ) ∈ X × U. Thus, (12.2) goes to 0. (c) Finally, assume µn → µ in total variation. If g is
bounded, (12.2) converges to 0, as in item (b). As a special case, assume that µn and µ admit densities hn and h,
respectively; then the pointwise convergence of hn to h implies the convergence of µn to µ in total variation by
Scheffé’s theorem.
(iv)Suppose now neither F nor the probability model of wt is known perfectly. It is assumed that wt admits a measure
µn and µn → µ weakly. For the function F we again have an approximating sequence {Fn }. If Fn (x, u, wn ) →
F (x, u, w) for all (x, u) ∈ X × U and for any wn → w, then the transition kernel Tn corresponding to the model
Fn converges weakly to the one of F , T : For any g ∈ Cb (X),
Z Z
lim g(x1 )Tn (dx1 |x0 , u0 ) = lim g(Fn (x0 , u0 , w))µn (dw)
n→∞ n→∞
Z Z
= g(F (x0 , u0 , w))µ(dw) = g(x1 )T (dx1 |x0 , u0 ).

(v) Let again {Fn } denote an approximating sequence for F and suppose now Fx0 ,u0 ,n (·) := Fn (x0 , u0 , ·) : W →
−1
X is invertible for all x0 , u0 ∈ X × U and F(x 0 ,u0 ),n
(·) denotes the inverse for fixed (x0 , u0 ). It is assumed
−1 −1
that F(x0 ,u0 ),n (x1 ) → Fx0 ,u0 (x1 ) pointwise for all (x0 , u0 ). Suppose further that the noise process wt admits a
continuous density fW (w). The Jacobian matrix, ∂x
∂w , is the matrix whose components are the partial derivatives
1

1 )i
of x1 , i.e. with x1 ∈ X ⊂ R and w ∈ W ⊂ R , it is an m × m matrix with components ∂(x
m m
∂wj , 1 ≤ i, j ≤ m
. If the Jacobian matrix of derivatives ∂x
∂w (w) is continuous in w and nonsingular for all w, then we have that the
1

density of the state variables can be written as

∂x1 −1 −1
fX1 ,n,(x0 ,u0 ) (x1 ) = fW (Fx−1
0 ,u0 ,n
(x1 )) (F (x1 )) ,
∂w x0 ,u0 ,n
∂x1 −1 −1
fX1 ,(x0 ,u0 ) (x1 ) = fW (Fx−1
0 ,u0
(x1 )) (F (x1 )) .
∂w x0 ,u0
With the above, fX1 ,n,(x0 ,u0 ) (x1 ) → fX1 ,(x0 ,u0 ) (x1 ) pointwise for all fixed (x0 , u0 ). Therefore, by Scheffé’s
theorem, the corresponding kernels Tn (·|x0 , u0 ) → T (·|x0 , u0 ) in total variation for all (x0 , u0 ).
(vi)These examples will be utilized in Section 12.5.1, where data-driven stochastic control problems will be considered
where estimated models are obtained through empirical measurements of the state action variables.

12.1.2 Summary

We now introduce the main assumptions that will be occasionally used for the technical results in the chapter.

Assumption 12.1.1 The following hold.


(a) The sequence of transition kernels Tn satisfies the following: {Tn (·|xn , un ), n ∈ N} converges weakly to T (·|x, u)
for any sequence {xn , un } ⊂ X × U and x, u ∈ X × U such that (xn , un ) → (x, u).
(b) The stochastic kernel T (·|x, u) is weakly continuous in (x, u).
(c) The sequence of stage-wise cost functions cn satisfies the following: cn (xn , un )
→ c(x, u) for any sequence {xn , un } ⊂ X × U and x, u ∈ X × U such that (xn , un ) → (x, u).
(d) The stage-wise cost function c(x, u) is non-negative, bounded, and continuous on X × U.
12.1 Introduction 295

(e) U is compact.

The following is Assumption 6.3.1(ii).

Assumption 12.1.2 The observation channel Q(·|x) is continuous in total variation i.e., if xn → x, then Q( · |xn ) →
Q( · |x) in total variation (only for partially observed models).

Assumption 12.1.3 The following hold.


(a) The sequence of transition kernels Tn satisfies the following: {Tn (·|x, un ), n ∈ N} converges setwise to T (·|x, u)
for any sequence {un } ⊂ U and x, u ∈ X × U such that un → u.
(b) The stochastic kernel T (·|x, u) is setwise continuous in u.
(c) The sequence of stage-wise cost functions cn satisfies the following: cn (x, un )
→ c(x, u) for any sequence {un } ⊂ U and x, u ∈ X × U such that un → u.
(d) The stage-wise cost function c(x, u) is non-negative, bounded, and continuous on U.
(e) U is compact.

Assumption 12.1.4 The following hold.


(a) The sequence of transition kernels Tn satisfies the following: ∥Tn (·|x, un ) − T (·|x, u)∥T V → 0 for any sequence
{un } ⊂ U and x, u ∈ X × U such that un → u.
(b) The stochastic kernel T (·|x, u) is continuous in total variation in u.
(c) The sequence of stage-wise cost functions cn satisfies the following: cn (x, un )
→ c(x, u) for any sequence {un } ⊂ U and x, u ∈ X × U such that un → u.
(d) The stage-wise cost function c(x, u) is non-negative, bounded, and continuous on U.
(e) U is compact.

In Sections 12.2.1 and 12.2.2 we study continuity (Problem P1) and robustness (Problem P2) for partially observed models.
In particular we show the following:
(a) Continuity and robustness do not hold in general under weak convergence of kernels (Theorem 12.2.1).
(b) Under Assumptions 12.1.1 and 12.1.2, continuity and robustness hold (Theorem 12.2.4, Theorem 12.2.8).
(c) Continuity and robustness do not hold in general under setwise convergence of the kernels (Theorem 12.2.5).
(d) Continuity and robustness do not hold in general under total variation convergence of the kernels (Example 12.1).
(e) Under Assumption 12.1.4, continuity and robustness hold (Theorem 12.2.6, Theorem 12.2.7).
In Section 12.3, we study continuity (Problem P1) and robustness (Problem P2) for fully observed models. In particular we
show the following
(a) Continuity and robustness do not hold in general under weak convergence of kernels (Theorem 12.3.1, Example
12.1).
(b) Under Assumption 12.1.1, continuity holds (Theorem 12.3.2), under Assumption 12.1.1, robustness holds if the
optimal policies for every initial point are identical (Theorem 12.3.3).
(c) Continuity and robustness do not hold in general under setwise convergence of the kernels (Theorem 12.3.4, Theo-
rem 12.3.6).
(d) Under Assumption 12.1.3, continuity holds (Theorem 12.3.5), and under Assumption 12.1.3, robustness holds if
the optimal policies for every initial point are identical (Theorem 12.3.7).
296 12 Robustness to Incorrect Models and Learning

(e) Continuity and robustness do not hold in general under total variation convergence of the kernels (Example 12.1).
(f) Under Assumption 12.1.4, continuity and robustness hold (subsection 12.3.3).
In Section 12.5, we study applications to empirical learning (in Section 12.5.1) where we establish the positive relevance
of Theorem 12.3.2, and then applications to finite model approximations under the perspective of robustness in in Section
12.5.2. Here, we restrict the analysis to the case with weakly continuous kernels.

12.2 Continuity and Robustness of Optimal Cost in Convergence of Models (POMDP Case)

12.2.1 Continuity of Optimal Cost in Convergence of Models (POMDP Case)

We now study continuity of the optimal discounted cost under the convergence of transition kernels and cost functions.

Weak Convergence

Absence of Continuity under Weak Convergence. The following shows that the optimal cost may not be continuous under
weak convergence of transition kernels.

Theorem 12.2.1 [186]. Let Tn → T weakly, then it is not necessarily true that Jβ∗ (c, Tn ) → Jβ∗ (c, T ) even when the
prior distributions are the same, the measurement channel Q is continuous in total variation, and c(x, u) is continuous and
bounded on X × U.

We prove the result with a counterexample [186]. Letting X = U = Y = [−1, 1] and c(x, u) = (x − u)2 , the observation
channel is chosen to be uniformly distributed over [-1,1], Q ∼ U ([−1, 1]), the initial distributions of the state variable are
chosen to be same as P ∼ δ1 , where δx (A) := 1{x∈A} for Borel A, and the transition kernels are:

1 1 1 1
T (·|x, u) = δ−1 (x)[ δ1 (·) + δ−1 (·)] + δ1 (x)[ δ1 (·) + δ−1 (·)]
2 2 2 2
+ (1 − δ−1 (x))(1 − δ1 (x))δ0 (·)
1 1 1
Tn (·|x, u) = δ−1 (x)[ δ(1−1/n) (·) + δ(−1+1/n) (·)] + δ1 (x)[ δ(1−1/n) (·)
2 2 2
1
+ δ(−1+1/n) (·)] + (1 − δ−1 (x))(1 − δ1 (x))δ0 (·).
2

It can be seen that Tn → T weakly according to Definition 12.1.1(i). Note that the cost function is continuous, and the
measurement channel is continuous in total variation. The optimal discounted costs can be found as
∞ ∞
X X β
Jβ∗ (c, T ) = EPT [β k Xk2 ] = βk =
1−β
k=1 k=1

X 1 1 1 1
Jβ∗ (c, Tn ) = EPTn [β k Xk2 ] = β[ (1 − )2 + (−1 + )2 ].
2 n 2 n
k=1

β
Then we have Jβ∗ (c, Tn ) → β ̸= 1−β .

A Sufficient Condition for Continuity under Weak Convergence

In the following, we will establish and utilize some regularity properties for the optimal cost with respect to the convergence
of transition kernels.
12.2 Continuity and Robustness of Optimal Cost in Convergence of Models (POMDP Case) 297

Assumption 12.2.1 (a) The stochastic kernel T (·|x, u) is weakly continuous in (x, u), i.e. if (xn , un ) → (x, u), then
T (·|xn , un ) → T (·|x, u) weakly.
(b) The observation channel Q(·|x) is continuous in total variation, i.e., if xn → x, then Q( · |xn ) → Q( · |x) in total
variation.
(c) The stage-wise cost function c(x, u) is non-negative, bounded and continuous on X × U
(d) U is compact.

As we have see in Chapter 6, any POMDP can be reduced to a (completely observable) MDP, whose states are the pos-
terior state distributions or beliefs of the observer; that is, the state at time t is Zt ( · ) := Pr{Xt ∈ · |Y0 , . . . , Yt , U0 , . . . ,
Ut−1 } ∈ P(X). We call this equivalent MDP the belief-MDP . The belief-MDP has state space Z = P(X) and action
space U. Under the topology of weak convergence, since X is a Borel space, Z is metrizable with the Prokhorov metric
which makes Z a Borel space [249]. The transition probability η (6.3) of the belief-MDP was earlier constructed through
non-linear filtering equations.
TheR one-stage cost function c of the belief-MDP is given by c̃(z, u)
:= X c(x, u)z(dx). Under the regularity of the belief-MDP, we have shown that the discounted cost optimality opera-
tor T : Cb (Z) → Cb (Z)

(T (f ))(z) = min(c̃(z, u) + βE[f (z1 )|z0 = z, u0 = u]) (12.3)


u

is a contraction from Cb (Z) to itself under the supremum norm. As a result, there exists a fixed point, the value function,
and an optimal control policy exists. In view of this existence result, in the following we will consider optimal policies.
The following result is key to proving the main result of this section whose detailed analysis can be found in [186].

Theorem 12.2.2 Suppose we have a uniformly bounded family of functions {fnγ : X → R, γ ∈ Γ, n > 0} such that
∥fnγ ∥∞ < C for all γ ∈ Γ and for all n > 0 for some C < ∞.
Further suppose we have another uniformly bounded family of functions {f γ : X → R, γ ∈ Γ } such that ∥f γ ∥∞ < C for
all γ ∈ Γ for some C < ∞. Under the following assumptions,
(i) For any xn → x

sup fnγ (xn ) − f γ (x) → 0, sup f γ (xn ) − f γ (x) → 0,


γ∈Γ γ∈Γ

(ii) supγ ρ(µγn , µγ ) → 0 where ρ is some metric for the weak convergence topology,
we have
Z Z
sup fnγ (x)µγn (dx) − f γ (x)µγ (dx) → 0.
γ∈Γ

Theorem 12.2.3 Under Assumptions 12.1.1 and 12.1.2,

sup |Jβ (cn , Tn , γ) − Jβ (c, T , γ)| → 0.


γ∈Γ

Proof sketch.

sup |Jβ (cn , Tn , γ) − Jβ (c, T , γ)|


γ∈Γ
∞  h 
X
t Tn
i T
h i
= sup β EP cn Xt , γ(Y[0,t] ) − EP c Xt , γ(Y[0,t] )
γ∈Γ t=0
298 12 Robustness to Incorrect Models and Learning
∞ h
X i h i
≤ β t sup EPTn cn Xt , γ(Y[0,t] ) − EPT c Xt , γ(Y[0,t] ) .
t=0 γ∈Γ

Recall that an admissible policy γ is a sequence of control functions {γt , t ∈ Z+ }. In the last step above, we make a slight
abuse of notation; the sup at the first step is over all sequence of control functions {γt , t ∈ Z+ } whereas the sup at the last
step is over all sequence of control functions {γt′ , t′ ≤ t}, but we will use the same notation, γ, in the rest of the proof.
P∞
For any ϵ > 0, we choose a K < ∞ such that t=K+1 β k 2∥c∥∞ ≤ ϵ/2. For the chosen K, we choose an N < ∞ such
that
h i h i
sup EPTn cn Xt , γ(Y[0,t] ) − EPT c Xt , γ(Y[0,t] ) ≤ ϵ/2K
γ∈Γ

for all t ≤ K and for all n > N . We note that in [186] a fixed c function was considered, but by considering the additional
term    
sup EPTn cn Xt , γ(Y[0,t] ) − EPT cn Xt , γ(Y[0,t] )
 
γ∈Γ
R R
and noting that supγ | Q(dy|xn )cn (xn , γ(y)) − Q(dy|x)c(x, γ(y))| → 0, for every xn → x, by a generalized domi-
nated convergence theorem as Q is continuous in total variation, a triangle inequality argument shows that the same result
applies. This follows from a generalized dominated convergence theorem as stated in Theorem 12.2.2 whose detailed
analysis can be found in [186]. Thus, supγ∈Γ Jβ (cn , Tn , γ) − Jβ (c, T , γ) → 0 as n → ∞. ⋄

Theorem 12.2.4 [182, 186] Suppose the conditions of Theorem 12.2.3 hold. Then

lim |Jβ∗ (cn , Tn ) − Jβ∗ (c, T )| = 0.


n→∞

Proof sketch. We start with the following bound:

|Jβ∗ (cn , Tn ) − Jβ∗ (c, T )| (12.4)


 
≤ max Jβ (cn , Tn , γ ∗ ) − Jβ (c, T , γ ∗ ), Jβ (c, T , γn∗ ) − Jβ (cn , Tn , γn∗ ) ,

where γ ∗ and γn∗ are the optimal policies, respectively, for T and Tn . Both terms go to 0 by Theorem 12.2.3. ⋄

Absence of Continuity under Setwise Convergence

We now show that continuity of optimal costs may fail under the setwise convergence of transition kernels. Theorem 12.3.4
in the next section establishes this result for fully observed models, which serves as a proof for this setup also.

Theorem 12.2.5 [182, 186] Let Tn → T setwise. Then, it is not true in general that

Jβ∗ (c, Tn )

→ Jβ∗ (c, T ), even when X, Y, and U are compact and c(x, u) is continuous and bounded in X × U.

Continuity under Total Variation

Theorem 12.2.6 [182, 186] Under Assumption 12.1.4, Jβ∗ (cn , Tn ) → Jβ∗ (c, T ).

Proof sketch. We start with the following bound:


12.2 Continuity and Robustness of Optimal Cost in Convergence of Models (POMDP Case) 299

|Jβ∗ (cn , Tn ) − Jβ∗ (c, T )| ≤ max Jβ (cn , Tn , γ ∗ ) − Jβ (c, T , γ ∗ ), Jβ (cn , Tn , γn∗ )

− Jβ (c, T , γn∗ ) ,

where γ ∗ and γn∗ are the optimal policies, respectively, for T and Tn .
We now study the following:

sup |Jβ (cn , Tn , γ) − Jβ (c, T , γ)|


γ∈Γ
∞  h 
X i h i
= sup β t
EPTn cn T
Xt , γ(Y[0,t] ) − EP c Xt , γ(Y[0,t] )
γ∈Γ t=0
∞ h
X i h i
≤ β t sup EPTn cn Xt , γ(Y[0,t] ) − EPT c Xt , γ(Y[0,t] ) .
t=0 γ∈Γ

It can be shown that ( [186])


h i h i
sup EPTn cn Xt , γ(Y[0,t] ) − EPT c Xt , γ(Y[0,t] ) → 0. (12.5)
γ∈Γ

This was shown in [186] for fixed c. The extension to varying cn follows from a triangle inequality step with the assumption
that Tn (·|x, un ) → T (·|x, u) setwise, and cn (x, un ) → c(x, u) for any un → u. Therefore, using identical steps as in the
proof of Theorem 12.2.3 we have supγ∈Γ Jβ (cn , Tn , γ) − Jβ (c, T , γ) → 0. ⋄

12.2.2 Robustness to Incorrect Models (POMDP Case)

Here, we consider the robustness problem P2: Suppose we design an optimal policy, γn∗ , for a transition kernel, Tn and a
cost function cn , assuming they are the correct model and apply the policy to the true model whose transition kernel is T
and whose cost function is c. We study the robustness of the sub-optimal policy γn∗ .

Total Variation

The next theorem gives an asymptotic robustness result.

Theorem 12.2.7 Under Assumption 12.1.4

|Jβ (cn , T , γn∗ ) − Jβ∗ (c, T )| → 0,

where γn∗ is the optimal policy designed for the kernel Tn .

Proof sketch. We write the following:

|Jβ (c, T , γn∗ ) − Jβ∗ (c, T )| ≤ |Jβ (c, T , γn∗ ) − Jβ∗ (cn , Tn )| + |Jβ∗ (cn , Tn ) − Jβ∗ (c, T )|.

Both terms can be shown to go to 0 using (12.5). ⋄


300 12 Robustness to Incorrect Models and Learning

Setwise Convergence

Theorem 12.3.6 in the next section establishes the lack of robustness under the setwise convergence of kernels. As we note
later, a fully observed system can be viewed as a partially observed system with the measurement being the state itself,
(see (12.6)).

Weak Convergence

Theorem 12.2.8 [182,186] Under Assumptions 12.1.1 and 12.1.2, |Jβ (c, T , γn∗ )−Jβ∗ (c, T )| → 0, where γn∗ is the optimal
policy designed for the transition kernel Tn .

Proof sketch. We write

|Jβ (c, T , γn∗ ) − Jβ∗ (c, T )| ≤|Jβ (c, T , γn∗ ) − Jβ (cn , Tn , γn∗ )| + |Jβ (cn , Tn , γn∗ )
− Jβ (T , γ ∗ )|.

The first term goes to 0 by Theorem 12.2.3. For the second term we use Theorem 12.2.4. ⋄

12.3 Continuity and Robustness in the Fully Observed Case

In this section, we consider the fully observed case where the controller has direct access to the state variables. We present
the results for this case separately, since here we cannot utilize the regularity properties of measurement channels which
allows for stronger continuity and robustness results. Under measurable selection conditions due to weak or strong (setwise)
continuity of transition kernels [164, Section 3.3], for infinite horizon discounted cost problems optimal policies can be
selected from those which are stationary and deterministic. Therefore we will restrict the policies to be stationary and
deterministic so that Ut = γ(Xt ) for some measurable function γ. Notice also that fully observed models can be viewed
as partially observed with the measurement channel thought to be

Q(·|x) = δx (·), (12.6)

which is only weakly continuous, thus it does not satisfy Assumption 12.1.2.

12.3.1 Weak Convergence

Absence of Continuity under Weak Convergence.

We start with a negative result.

Theorem 12.3.1 [182, 186] For Tn → T weakly, it is not necessarily true that Jβ∗ (c, Tn ) → Jβ∗ (c, T ) even when the prior
distributions are the same and c(x, u) is continuous and bounded in X × U.

Proof. We prove the result with a counterexample, similar to the model used in the proof of Theorem 12.2.1 Letting
X = [−1, 1], U = {−1, 1} and c(x, u) = (x − u)2 , the initial distributions are given by P ∼ δ1 , that is, X0 = 1, and the
transition kernels are
1 1 1 1
T (·|x, u) =δ−1 (x)[ δ1 (·) + δ−1 (·)] + δ1 (x)[ δ1 (·) + δ−1 (·)]
2 2 2 2
+ (1 − δ−1 (x))(1 − δ1 (x))δ0 (·),
12.3 Continuity and Robustness in the Fully Observed Case 301

1 1 1
Tn (·|x, u) =δ−1 (x)[ δ(1−1/n) (·) + δ(−1+1/n) (·)] + δ1 (x)[ δ(1−1/n) (·)
2 2 2
1
+ δ(−1+1/n) (·)] + (1 − δ−1 (x))(1 − δ1 (x))δ0 (·).
2
It can be seen that Tn → T weakly according to Definition 12.1.1(i). Under this setup we can calculate the optimal costs
as follows:

1 X 1 β2
Jβ∗ (c, Tn ) = + β k
= + ,
n2 n2 1−β
k=2

and Jβ∗ (c, T ) = 0. Thus, continuity does not hold. ⋄


We now present another counter example emphasizing the importance of continuous convergence in the actions. The
following counter example shows that without the continuous convergence and regularity assumptions on the kernel T ,
continuity fails even when Tn (·|x, u) → T (·|x, u) pointwise (for x, u) in total variation (also setwise and weakly) and
even when the cost function c(x, u) is continuous and bounded. Notice that this example also holds for both setwise and
weak convergence.

Example 12.1. Assume that the kernels are given by

Tn (·|x, u) ∼ U ([un , 1 + un ]),


(
U ([0, 1]) if u ̸= 1,
T (·|x, u) ∼
U ([1, 2]) if u = 1,

where U = [0, 1] and X = R. We note first that Tn (·|x, u) → T (·|x, u) in total variation for every fixed x and u.
The cost function is in the following form:

if x ≤ 1e ,

2

x− 1e

2 − if 1e < x ≤ 0.1 + 1e ,


0.1

c(x, u) = 1 if 0.1 + 1e < x ≤ 1 + 1
e − 0.1,
1+ 1e −x

1
2− − 0.1 < x ≤ 1 + 1e ,



 0.1 if 1 + e

1
2 if 1 + < x.

e

Notice that c(x, u) is a continuous function.


With this setup, γ ∗ (x) = 0 is an optimal policy for T since on the [0, 1] interval the induced cost is less than the cost
induced on the [1, 2] interval. The cost under this policy is
∞    
X 1 0.3 1 1 1
Jβ∗ (c, T )= t
β 2× + + 0.9 − = 1.05 + .
t=0
e 2 e 1−β e

1 1 1 1
For Tn , γn∗ (x) = e− n is an optimal policy for every n as e− n ×n = e and thus the state is distributed between e <x≤
1 + 1e in which interval the cost is the least. Hence, we can write
∞    
X 1.1 1 1
lim Jβ (c, Tn , γn∗ ) = β t 0.3 + 1 − 0.2 = ̸ = 1.05 +
n→∞
t=0
1−β 1−β e
= Jβ∗ (c, T ).


302 12 Robustness to Incorrect Models and Learning

A Sufficient Condition for Continuity under Weak Convergence.

We will now establish that if the kernels and the model components have some further regularity, continuity does hold.
The assumptions of the following result are the same as the assumptions for the partially observed case (Theorem 12.2.4)
except for the assumption on the measurement channel Q.

Theorem 12.3.2 [182, 186] Under Assumption 12.1.1, Jβ (cn , Tn , γn∗ ) → Jβ (c, T , γ ∗ ) for any initial state x0 , as n → ∞.

Proof. We will use the successive approximations for an inductive argument.


Recall the discounted cost optimality operator T : Cb (Z) → Cb (Z) from (12.3)
 
(T (v))(x) = inf c(x, u) + βE[v(x1 )|x0 = x, u0 = u] ,
u∈U

which is a contraction from Cb (X) to itself under the supremum norm and has a fixed point, the value function. For the
kernel T , we will denote the approximation functions by

v k (x) = T (v k−1 )(x),

and for the kernel Tn we will use vnk (x) to denote the approximation functions, notice that the operator T also depends on
n for the model Tn , but we will continue using it as T in what follows.
We wish to show that the approximation functions for Tn continuously converge to the ones for T . Then, for the first step
of the induction we have

v 1 (x) = c(x, u∗ ), vn1 (xn ) = cn (xn , u∗n ),

and thus we can write,

|v 1 (x) − vn1 (xn )| ≤ sup c(x, u) − cn (xn , u)


u∈U

since cn (xn , un ) → c(x, u) for all (xn , un ) → (x, u) and the action space, U, is compact, the first step of the induction
holds, i.e. limn→∞ |v 1 (x) − vn1 (xn )| = 0.
For the k th step we have
Z
v k (x) = T (v k−1 )(x) = inf c(x, u) + β v k−1 (x1 )T (dx1 |x, u) ,
 
u X
Z
vn (xn ) = T (vn )(xn ) = inf cn (xn , u) + β vnk−1 (x1 )Tn (dx1 |xn , u) .
k k−1
 
u X

Note that the assumptions of the theorem satisfy the measurable selection criteria and hence we can choose minimizing
selectors ( [164, Section 3.3]). If we denote the selectors by u∗ and u∗n , we can write

|v k (x) − vnk (xn )|



≤ max |c(x, u∗ ) − cn (xn , u∗ )|
Z Z 
+ β| v k−1 (x1 )T (dx1 |x, u∗ ) − vnk−1 (x1 )Tn (dx1 |xn , u∗ )| ,
X X

∗ ∗
|c(x, un ) − cn (xn , un )|
Z Z 
+ β| v k−1 (x1 )T (dx1 |x, u∗n ) − vnk−1 (x1 )Tn (dx1 |xn , u∗n )| .
X X
12.3 Continuity and Robustness in the Fully Observed Case 303

Hence, we can write

|v k (x) − vnk (xn )| (12.7)



≤ sup |c(x, u) − cn (xn , u)|
u∈U
Z Z 
+ β| v k−1 (x1 )T (dx1 |x, u) − vnk−1 (x1 )Tn (dx1 |xn , u)| ,
X X

above, the first term goes to 0 as cn (xn , un ) → c(x, u) for all (xn , un ) → (x, u) and the action space, U, is compact. For
the second term we write,
Z Z
sup | v k−1 (x1 )T (dx1 |x, u) − vnk−1 (x1 )Tn (dx1 |xn , u)|
u∈U
ZX X

k−1 1
(x ) − vn (x ) Tn (dx1 |xn , u)|
k−1 1

≤ sup | v
u∈U X
Z Z
+ sup | v k−1 (x1 )T (dx1 |x, u) − v k−1 (x1 )Tn (dx1 |xn , u)|
u∈U X X

above, for the first term, by the induction argument for any x1n → x1 , v k−1 (x1 )−vnk−1 (x1n ) → 0 (i.e., we have continuous
convergence).
We also have that Tn (·|xn , u) → T (·|x, u) weakly uniformly over u ∈ U as U is compact. Therefore, using Theorem
12.2.2 the first term goes to 0. For the second term we again use that Tn (·|xn , u) converges weakly to T (·|x, u) uniformly
over u ∈ U. With an almost identical induction argument it can also be shown that v k−1 (x1 ) is continuous in x1 , thus the
second term also goes to 0.
So far, we have showed that for any k ∈ N, limn→∞ vnk (xn ) − v k (x) = 0 for any xn → x, in particular it is also true
that limn→∞ vnk (x) − v k (x) = 0 for any x.
As we have stated earlier, it can be shown that the approximation operator, T is a contractive operator under supremum
norm with modulus β and it converges to a fixed point which is the value function. Thus, we have

βk βk
Jβ (c, T , γ ∗ ) − v k (x) ≤ ∥c∥∞ , Jβ∗ (cn , Tn , γn∗ ) − vnk (x) ≤ ∥c∥∞ . (12.8)
1−β 1−β

Combining the results,

|Jβ (cn , Tn , γn∗ ) − |Jβ (c, T , γ ∗ )| ≤|Jβ (cn , Tn , γn∗ ) − vnk (x)| + |vnk (x) − v k (x)|
+ |Jβ (c, T , γ ∗ ) − v k (x)|.

Note that the first and the last term can be made arbitrarily small since (12.8) holds for all k ∈ N; the second term goes to
0 with an inductive argument for all k ∈ N. ⋄

A Sufficient Condition for Robustness under Weak Convergence.

We now present a result that establishes robustness if the optimal policies for every initial point are identical. That is, for
every n, γn∗ is optimal for every x0 ∈ X (under the model Tn ). A sufficient condition for this property is that γn∗ solves the
discounted cost optimality equation (DCOE) for every initial point.
A policy γ ∗ ∈ Γ solves the discounted cost optimality equation and is optimal if it satisfies
Z
Jβ∗ (c, T , x) = c(x, γ ∗ (x)) + β Jβ∗ (c, T , x1 )T (dx1 |x, γ ∗ (x)).
304 12 Robustness to Incorrect Models and Learning

Thus, a policy is optimal for every initial point if it satisfies the DCOE for all initial points x ∈ X. The following generalizes
[186].

Theorem 12.3.3 Under Assumption 12.1.1, Jβ (cn , T , γn∗ ) → Jβ (c, T , γ ∗ ) for any initial point x0 if γn∗ is optimal for any
initial point for the kernel Tn and for the stage-wise cost function cn .

Remark 12.2. For the partially observed case, the proof approach we use makes use of policy exchange (e.g. (12.4)) and
for this approach the total variation continuity of channel Q(·|x) is a key step to deal with the uniform convergence over
policies. As we stated before, the channel for fully observed models can be considered in the form of (12.6) which is only
weakly continuous and not continuous in total variation. Thus, obtaining a result uniformly over all policies may not be
possible. However, for the fully observed models we can reach continuity and robustness (Theorem 12.3.2, Theorem 12.3.3)
using a value iteration approach. With this approach, instead of exchanging policies and analyzing uniform convergence
over all policies, we can exchange control actions (e.g. (12.7)) and analyze uniform convergence over the action space U
by using the discounted optimality operator (12.3). Hence, we are only able to show convergence over optimal policies for
the fully observed case, i.e. Jβ (cn , Tn , γn∗ ) → Jβ (c, T , γ ∗ ) or Jβ (c, T , γn∗ ) → Jβ (c, T , γ ∗ ) where γn∗ and γ ∗ are optimal
policies, whereas, for partially observed models, regularity of the channel allows us to show convergence over any sequence
of policies, i.e. supγ∈Γ |Jβ (cn , Tn , γ) − Jβ (c, T , γ)| → 0.

Remark 12.3. As we have discussed in subsection 12.2.1, a partially observed model can be reduced to a fully observed
process where the state process (beliefs) becomes probability measure valued. Consider the partially observed models with
transition kernels Tn and T (with a channel Q) and their corresponding fully observed transition kernels ηn and η: following
the discussions and techniques in [127] and [183], one can show that ηn and η satisfy the conditions of Theorem 12.3.3
and Theorem 12.3.2 that is ηn (·|zn , un ) → η(·|z, u) for any (zn , un ) → (z, u) under the following set of assumptions
– Tn (·|xn , un ) → T (·|x, u) for any (xn , un ) → (x, u),
– Q(·|x) is continuous on total variation in x.
We remark that these conditions also agree with the conditions presented for continuity and robustness of the partially
observed models (Theorem 12.2.4 and Theorem 12.2.8).

12.3.2 Setwise Convergence

Absence of Continuity under Setwise Convergence.

We give a negative result similar to Theorem 12.2.5, via Example 12.1:

Theorem 12.3.4 Letting Tn → T setwise, then it is not necessarily true that Jβ∗ (c, Tn ) → Jβ∗ (c, T ) even when c(x, u) is
continuous and bounded in X × U.

A Sufficient Condition for Continuity under Setwise Convergence.

Theorem 12.3.5 Under Assumption 12.1.3 Jβ (cn , Tn , γn∗ ) → Jβ (c, T , γ ∗ ), for any initial state x0 , as n → ∞.

Proof. We use the same value iteration technique that we used to prove Theorem 12.3.2. See [186]. ⋄

Absence of Robustness under Setwise Convergence.

Now, we give a result showing that even if the continuity holds under the setwise convergence of the kernels, the robustness
may not be satisfied (see [186, Theorem 4.7]).
12.4 The Average Cost Case 305

Theorem 12.3.6 Supposing Tn (·|xn , un ) → T (·|x, u) setwise for every x ∈ X and u ∈ U and (xn , un ) → (x, u), then it
is not true in general that Jβ (c, T , γn∗ ) → Jβ (c, T , γ ∗ ), even when X and U are compact and c(x, u) is continuous and
bounded in X × U.

A Sufficient Condition for Robustness under Setwise Convergence.

We now present a similar result to Theorem 12.3.3 that is we show that under the conditions of Theorem 12.3.5, if further
for every n, γn∗ is optimal for every x0 ∈ X (under the model Tn ) then robustness holds under setwise convergence.

Theorem 12.3.7 Supposing Assumption 12.1.3 holds, if further we have that for every n, γn∗ is optimal for every x0 ∈ X
(under the model Tn ) then Jβ (c, T , γn∗ ) → Jβ (c, T , γ ∗ ).

12.3.3 Total Variation

The continuity result in Theorem 12.2.6 and the robustness result in Theorem 12.2.7 apply to this case since the fully
observed model may be viewed as a partially observed model with the measurement channel Q given in (12.6).

12.4 The Average Cost Case

The results above also apply to the average cost setup by adding an ergodicity condition, as we have seen in Chapter 7.
In the following we will denote the set of all stationary policies by Γs . For the transitions under some stationary policy γ,
we will use the following notation: T (·|x, γ) := T (·|x, γ(x)).
We also define the t-step transition kernel T t (·|x, γ) in an iterative fashion as follows:
Z
T (·|x, γ) := T (·|xt−1 , γ)T t−1 (dxt−1 |x, γ),
t

where T 1 (·|x, γ) = T (·|x, γ).


We will use the following ergodicity condition for some of the results.

Assumption 12.4.1 For every stationary policy γ, the transition kernels T and Tn lead to positive Harris recurrent chains
and in particular admit invariant measures πγ and πγn , and for these invariant measures uniformly for every initial point
x ∈ X we have:

lim sup ∥T t (·|x, γ) − πγ (·)∥T V = 0


t→∞ γ∈Γs

lim sup sup ∥Tnt (·|x, γ) − πγn (·)∥T V = 0.


t→∞ n γ∈Γs

We have seen the following earlier in Chapter 7, repeated in a concise form for reader’s convenience:
For our continuity and robustness results, it will be instrumental to work with stationary policies. This will be without any
loss under mild conditions to be presented in this subsection. An approach for average cost problems is to make use of
average cost optimality equation (ACOE). To work with ACOE one usually needs contraction properties of the transition
kernel. The following result provides further alternative sufficient conditions on existence of optimal policies (which turn
out to be stationary) for infinite horizon average cost problems.

Assumption 12.4.2 The following hold.


306 12 Robustness to Incorrect Models and Learning

(A) Assumption 7.2.2 holds,


(B) The action space U is compact,
(C) c(x, u) is bounded and continuous in (x, u),
(C’)c(x, u) is bounded and continuous in u for every fixed x,
(D) T (·|x, u) is weakly continuous in (x, u),
(D’)T (·|x, u) is setwise continuous in u for every x.

Proposition 12.4. Suppose Assumption 12.4.2 A, B, and, either C and D, or C’ and D’, hold. Then J∞ (T , γ) admits an
optimal stationary policy. ⋄

12.4.1 Approximation by finite horizon cost

We denote the t-step finite horizon cost function under a stationary policy γ and a transition model T by Jt (T , γ) and the
corresponding optimal cost is denoted by Jt∗ (T ):
t−1
X
Jt (T , γ) = EγT [c(Xi , Ui )]
i=0
Jt∗ (T ) = inf Jt (T , γ).
γ∈Γ

The following result shows that the infinite horizon average cost induced by a stationary policy can be approximated by a
finite cost under the same stationary policies with proper ergodicity conditions.

Lemma 12.5. [181] Under Assumption 12.4.1, if the cost function c is bounded then for every initial state we have

Jt (T , γ)
sup − J∞ (T , γ) → 0,
γ∈Γ t
Jt (Tn , γ)
sup sup − J∞ (Tn , γ) → 0.
γ∈Γ n t

c(x, γ(x))π γ (dx). Thus, we can write


R
Proof. We have that J∞ (T , γ) =

Jt (T , γ)
− J∞ (T , γ)
t
t−1 Z
1X T
= Eγ [c(Xi , Ui )] − c(x, γ(x))π γ (dx)
t i=0
t−1 Z Z
1X
≤ c(xi , γ(xi ))T (dxi |x0 , γ) − c(x, γ(x))π γ (dx)
i
t i=0
t−1
1X
≤ ∥c∥∞ ∥T i (·|x0 , γ) − π γ ∥T V .
t i=0

We now fix an ϵ > 0 and choose a tϵ < ∞ such that ∥T i (·|x0 , γ) − π γ ∥T V < ϵ for all i > tϵ . We also choose another Tϵ
with 2ttϵ < ϵ for all t > Tϵ . With this setup, we have
t−1
1X i
∥T (·|x0 , γ) − π γ ∥T V
t i=0
12.4 The Average Cost Case 307
tϵ −1 t
1 X 1X
≤ ∥Tγi (·|x0 ) − π γ ∥T V + ∥Tγi (·|x0 ) − π γ ∥T V
t i=0 t i=t
ϵ

2tϵ
≤ + ϵ ≤ 2ϵ, ∀t > Tϵ .
t
We have shown that for any fixed ϵ > 0, we can choose a Tϵ < ∞, independent of γ, such that

Jt (T , γ)
− J∞ (T , γ) < ϵ, ∀t > Tϵ .
t

Hence the result is complete for T .


For Tn the result follows from the same steps since we can again choose such tϵ and Tϵ due to the uniformity over n and γ
in Assumption 12.4.1. ⋄
The next result from [159, Corollary 4.11] shows that the optimal infinite horizon cost can be approximated by an optimal
finite horizon cost induced by the same transition kernel.

Lemma 12.6. Suppose the cost function c is bounded and either Assumption 12.4.2 A, B, C, D or Assumption 12.4.2 A, B,
C’, D’ hold (for T and Tn ). Then, we have

∗ Jt∗ (T )
lim J∞ (T ) − → 0,
t→∞ t
∗ J ∗ (Tn )
lim sup J∞ (Tn ) − t → 0.
t→∞ n t

12.4.2 Continuity under the convergence of transition kernels

∗ ∗
Theorem 12.7. [181] We have that |J∞ (Tn ) − J∞ (T )| → 0, under
c1. Assumption 12.4.2 A, B, C and D if Tn (·|xn , un ) → T (·|x, u) weakly for any (xn , un ) → (x, u).
c2. Assumption 12.4.2 A, B, C ′ and D′ if Tn (·|x, un ) → T (·|x, u) setwise for any un → u for every fixed x.

Proof. We use the following bound:


∗ ∗
|J∞ (Tn ) − J∞ (T )|
∗ Jt∗ (Tn )
≤ J∞ (Tn ) −
t
∗ ∗
J (Tn ) Jt (T ) J ∗ (T ) ∗
+ t − + t − J∞ (T ) .
t t t

The first and the last terms above can be made arbitrarily small by choosing t large enough uniformly over n using Lemma
12.6 under suitable assumptions. For the second term, we can use continuity results for finite time problems for the fixed
t as the assumptions cover the requirements studied earlier: see the proofs of Theorem 12.3.2 (for weak convergence) and
Theorem 12.3.5 (for setwise convergence). ⋄

12.4.3 Robustness to Incorrect Controlled Transition Kernel Models

In this section, we investigate robustness for infinite horizon average cost problems. We first restate the problem: Consider
an MDP with transition kernel Tn , and assume that an optimal control policy for this MDP under the average cost criterion
is γn∗ , that is
308 12 Robustness to Incorrect Models and Learning

inf J∞ (Tn , γ) = J∞ (Tn , γn∗ ).


γ∈Γ


Now, consider another MDP with transition kernel T whose optimal cost is denoted by J∞ (T ). If the controller does not
know the true transition kernel T and calculates an optimal policy assuming the transition kernel is Tn , then the incurred
cost by this policy is J∞ (T , γn∗ ). The focus of this section is to find sufficient conditions such that as Tn → T ,

J∞ (T , γn∗ ) → J∞ (T , γ ∗ ).

Suppose that the MDP with kernel Tn admits two different optimal policies γn1 and γn2 . Although, the cost incurred by these
policies under the kernel Tn are the same, under the kernel T they may have different cost values. That is, even though we
have that

J∞ (Tn , γn1 ) = J∞ (Tn , γn2 ) = J∞



(Tn ),

we may have J∞ (T , γn1 ) ̸= J∞ (T , γn2 ). An example is as follows: Consider a system with state space X = [−1, 1], control
action space U = {−1, 0, 1}, the cost function c(x, u) = (x − u)2 and the transition models are given as
1 1
Tn (·|x, u) = δ1 (·) + δ−1 (·)
2 2
T (·|x, u) = δ0 (·).

Notice that two optimal policies for Tn are



1 if x = 1,
 (
1 1 if x ≥ 0,
γn (x) = −1 if x = −1, γn2 (x) =
 −1 if x < 0.
0 else.

However, if the initial point is x0 = 0, we have that J∞ (T , γn1 ) = 0 ̸= 1 = J∞ (T , γn2 ).


In what follows, we show that under total variation convergence of Tn → T , this issue does not cause a problem so that we
have J∞ (T , γn∗ ) → J∞ (T , γ ∗ ) for any stationary optimal policy γn∗ . However, under weak or setwise convergence of the
transition models, we establish the same result under some particularly constructed optimal polices γn∗ , namely we focus
on the policies that solve the average cost optimality equation (ACOE) (in analogy with the corresponding results under
the discounted cost criterion: Theorem 12.3.3 under weak continuity, Theorem 12.3.7 for setwise continuity).
The following summarize some of the relevant findings in Section 7.2.

Proposition 12.8. [159] Suppose the cost function c is bounded. Under Assumption 12.4.1, there exists a β < 1 such that
the following holds:

(i) sp T v − T w ≤ β sp(u − w), for any v, w ∈ B(X) where T is the operator defined in (7.21).
(ii) Since T is a contraction under the span norm, it admits a fixed point v ∗ ∈ B(X) such that
 Z 
∗ ∗ ∗
j + v (x) = inf c(x, u) + v (y)T (dy|x, u) ,
u∈U X

for some constant j ∗ .


(iii)For any initial point x0 ∈ X, the constant j ∗ defined in (ii) is the optimal infinite horizon average cost for the
kernel T , that is

j ∗ = J∞

(T , x0 ) = inf J∞ (T , γ, x0 )
γ∈Γ

for every x0 ∈ X.
12.4 The Average Cost Case 309

(iv) If there exists a policy γ ∗ ∈ Γ satisfying the ACOE, then this stationary policy is an optimal policy for the average
infinite horizon cost problem; that is, if γ ∗ satisfies
Z
∗ ∗ ∗
j + v (x) = c(x, γ (x)) + v ∗ (y)T (dy|x, γ ∗ (x)),
X

then J∞ (T , γ ∗ , x0 ) = J∞

(T , x0 ).

Theorem 12.9. [181] We have that


J∞ (T , γn∗ , x) → J∞

(T , x)
for any x ∈ X, where γn∗ is the optimal policy for the transition kernel Tn that satisfies the ACOE, under Assumption
12.4.2 A, B, C and D if Tn (·|xn , un ) → T (·|x, u) weakly for any (xn , un ) → (x, u).

Proof. Consider the following two ACOEs for the kernels Tn and T with their fixed points vn∗ and v ∗ :
 Z 
∗ ∗ ∗
jn + vn (x) = inf c(x, u) + vn (y)Tn (dy|x, u) (12.9)
u∈U
 Z 
∗ ∗ ∗
j + v (x) = inf c(x, u) + v (y)T (dy|x, u) (12.10)
u∈U

We now show that, for all xn → x,

vn∗ (xn ) − v ∗ (x) → c (12.11)

for some constant c with |c| < ∞. To show this, we first write

vn∗ (xn ) − v ∗ (x)


= vn∗ (xn ) − vnt (xn ) + vnt (xn ) − v t (x) + v t (x) − v ∗ (x)
  

where vnt and v t are the results of operator (7.21) applied to the 0-function, t times for kernels the Tn and T . Notice
that vnt and v t are the value functions for t-step cost problem and by the assumptions ( [186, Theorem 4.4]) we have
that |vnt (xn ) − v t (x)| → 0 for every fixed t. For the first and the last terms, we use the fact that the operator (7.21) is
a contraction under Assumption 7.2.2 for the span semi-norm and hence both terms go to some constants as t → ∞
uniformly for all n, that is vn∗ (xn ) − vnt (xn ) → c1 and v t (x) − v ∗ (x) → c2 for some |c1 |, |c2 | < ∞. Thus, we have that
(12.11) holds for some c < ∞.
Since U is compact, for every xn → x, γn (xn ) has a convergent subsequence which converges to say some u∗ ∈ U . If we
take the limit along this subsequence for (12.9), using the assumptions that Tn (·|xn , un ) → T (·|x, u) weakly, the fact that
limn→∞ vn∗ (xn ) − v ∗ (x) = c, and that jn∗ → j ∗ (continuity results from Theorem 12.7) we get


 
lim jn∗k + vn∗ k (xnk )
k
Z
= lim c(x, γn∗k (xnk )) + vn∗ k (y)Tnk (dy|xnk , γn∗k (xnk ))
k
Z
= j + v (x) + c = c(x, u ) + v ∗ (y)T (dy|x, u∗ ) + c.
∗ ∗ ∗

Therefore, u∗ satisfies the ACOE for the kernel T and thus, any convergent subsequence of γn∗ (xn ) is an optimal action
for x for the kernel T .
Now consider the following operator T̂n , for the kernel T and the policy γn∗ which is optimal for Tn
Z
T̂n v̂n (x) = c(x, γn∗ (x)) + v̂n (y)T (dy|x, γn∗ (x)). (12.12)
310 12 Robustness to Incorrect Models and Learning

One can show that this operator is also a contraction under span semi-norm and admits a fixed point v̂n∗ , such that
Z
ĵn + v̂n (x) = c(x, γn (x)) + v̂n∗ (y)T (dy|x, γn∗ (x))
∗ ∗

where ĵn = J∞ (T , γn∗ , x) for all x. Hence, we need to show that ĵn → j ∗ to complete the proof. To show this, in [181], it
has been proven that

lim v̂n∗ (xn ) − v ∗ (x) = ĉ, (12.13)


n→∞

for any xn → x for some constant ĉ < ∞.


Now, assume that limn ĵn ̸= j ∗ and that there exists a subsequence ĵnk and an ϵ > 0 such that |ĵnk − j ∗ | > ϵ for every k.
We will show that this cannot hold, by establishing the existence of a further subsequence ĵnkl which converges to j ∗ in
the following.
We first note that limn→∞ v̂n∗ (xn ) − v ∗ (x) = ĉ. Hence, Theorem D.3.1 yields that v̂n∗ k (y)T (dy|x, γn∗k (x)) →
R
R ∗ l l
v (y)T (dy|x, u∗ ) + ĉ where ĉ also satisfies v̂n∗ k (x) − v ∗ (x) → ĉ.

l

Therefore, taking the limit along this subsequence,

lim ĵnkl
l→∞
Z
= lim c(x, γn∗k (x)) + v̂n∗ k (y)T (dy|x, γn∗k (x)) − v̂n∗ k (x)
l→∞ l l l l
Z
= c(x, u∗ ) + v ∗ (y)T (dy|x, u∗ ) − v ∗ (x) = j ∗ .

This contradicts to |ĵnk − j ∗ | > ϵ, hence we conclude that ĵn → j ∗ . ⋄


We now obtain the same under setwise convergence, without proof.

Theorem 12.10. [181] We have that J∞ (T , γn∗ , x) → J∞ ∗


(T , x) for any x ∈ X, where γn∗ is the optimal policy for the
transition kernel Tn that satisfies the ACOE, under Assumption 12.4.2 A, B, C and D if Tn (·|x, un ) → T (·|x, u) setwise
for any un → u.

For total variation, a more direct result follows.

Theorem 12.11. [181] We have that |J∞ (T , γn∗ ) − J∞∗


(T )| → 0 for any stationary optimal policy γn∗ for Tn , under
′ ′
Assumption 12.4.2 A, B, C and D if Tn (·|x, un ) → T (·|x, u) in total variation for any un → u for every fixed x.

Proof. We write:

|J∞ (T , γn∗ ) − J∞

(T )|
∗ ∗
≤ |J∞ (Tn ) − J∞ (T )| + |J∞ (T , γn∗ ) − J∞

(Tn )|

the first term goes to by Theorem 12.7. For the second term we write

|J∞ (T , γn∗ ) − J∞

(Tn )|
Jt (Tn , γn∗ )
≤ J∞ (T , γn∗ ) −
t
Jt (Tn , γn∗ ) Jt (T , γn∗ ) Jt (T , γn∗ )
+ − + − J∞ (T , γn∗ )
t t t
12.5 Applications to Data-Driven Learning and Finite Model Approximations 311

The first and the last terms above again can be made arbitrarily small by choosing t large enough uniformly over n
using Lemma 12.5. For the second term we use [186, Section A.2] where it is shown that under the stated assumptions
supγ∈Γ |Jt (Tn , γ) − Jt (T , γ)| → 0. Hence the proof is complete. ⋄

12.5 Applications to Data-Driven Learning and Finite Model Approximations

12.5.1 Application of Robustness Results to Data-Driven Learning

In practice, one may estimate the kernel of a controlled Markov chain using empirical data.
Let us briefly review the basic case where an i.i.d. sequence of random variables is repeatedly observed, but its probability
measure is not known apriori. Let {(Xi ), i ∈ N} be an X-valued i.i.d. random variable sequence generated according to
Pndistribution µ. Defining for every (fixed) Borel B ⊂ X, and n ∈ N, the empirical occupation measures µn (B) =
some
1
n i=1 1{Xi ∈B} , one has µn (B) → µ(B) almost surely by the strong law of large numbers. It then follows that µn → µ
weakly with probability one ( [110], Theorem 11.4.1). However, µn does not converge to µ in total variation or setwise, in
general. On the other hand, if we know that µ admits a density, we can find estimators to estimate µ under total variation
[103, Chapter 3]. In the previous sections, we established robustness results under the convergence of transition kernels in
the topology of weak convergence and total variation. We build on these observations.

Corollary 12.12 (to Theorem 12.2.6 and Theorem 12.2.7 ). Suppose we are given the following dynamics for finite state
space, X, and finite action space, U,

xt+1 = f (xt , ut , wt ), yt = g(xt , vt )

where {wt } and {vt } are [Link] processes and the noise models are unknown. Suppose that there is an initial training
period so that under some policy, every x, u pair is visited infinitely often if training were to continue indefinitely, but that
the training ends at some finite time. Let us assume that, through this training, we empirically learn the transition dynamics
such that for every (fixed) Borel B ⊂ X, for every x ∈ X, u ∈ U and n ∈ N, the empirical occupation measures are
Pn
1{X ∈B,Xi−1 =x,Ui−1 =u}
Tn (B|x0 = x, u0 = u) = i=1 Pn i .
i=1 1{Xi−1 =x,Ui−1 =u}

Then we have that Jβ∗ (Tn ) → Jβ∗ (T ) and Jβ (T , γn∗ ) → Jβ∗ (T ), where γn∗ is the optimal policy designed for Tn . Since the
channel model g has no restrictions, this result also applies to the fully observed model setup by taking g(xt , vt ) = xt .

Proof. We have that Tn (·|x, u) → T (·|x, u) weakly for every x ∈ X, u ∈ U almost surely by law of large numbers. Since
the spaces are finite, we also have Tn (·|x, u) → T (·|x, u) under total variation. By Theorem 12.2.6 and Theorem 12.2.7,
the results follow. ⋄

The following holds for more general spaces.

Corollary 12.13 (to Theorems 12.2.8, 12.2.4, 12.3.2 and 12.3.3). Suppose we are given the following dynamics with state
space X and action space U,
xt+1 = f (xt , ut , wt ), yt = g(xt , vt ),
where {wt } and {vt } are [Link] processes and the noise models are unknown. Suppose that f (x, u, ·) : W → X is
invertible for all fixed (x, u) and f (x, u, w) is continuous and bounded on X×U×W. We construct the empirical measures
for the noise process wt such that for every (fixed) Borel B ⊂ W, and for every n ∈ N, the empirical occupation measures
are
n
1X
µn (B) = 1 −1 (12.14)
n i=1 {fxi−1 ,ui−1 (xi )∈B}
312 12 Robustness to Incorrect Models and Learning

where fx−1
i−1 ,ui−1
(xi ) denotes the inverse of f (xi−1 , ui−1 , w) : W → X for given (xi−1 , ui−1 ). Using the noise measure-
ments, we construct the empirical transition kernel estimates for any (x0 , u0 ) and Borel B as

Tn (B|x0 , u0 ) = µn (fx−1
0 ,u0
(B)).

(i) If the measurement channel (represented by the function g) is continuous in total variation then Jβ∗ (Tn ) → Jβ∗ (T )
and Jβ (T , γn∗ ) → Jβ∗ (T ), where γn∗ is the optimal policy designed for Tn for all initial points.
(ii) If the measurement channel is in the form g(xt , vt ) = xt (i.e. fully observed) then Jβ∗ (Tn ) → Jβ∗ (T ) and if further
for every n, γn∗ is optimal for every x0 ∈ X (under the model Tn ) then Jβ (T , γn∗ ) → Jβ∗ (T ).

Proof. We have µn → µ weakly with probability one where µ is the model. We claim that the transition kernels are such
that Tn (·|xn , un ) → T (·|x, u) weakly for any (xn , un ) → (x, u). To see that observe the following for h ∈ Cb (X)
Z Z
h(x1 )Tn (dx1 |xn , un ) − h(x1 )T (dx1 |x, u)
Z Z
= h(f (xn , un , w))µn (dw) − h(f (x, u, w))µ(dw) → 0,

where µn is the empirical measure for wt and µ is the true measure again. For the last step, we used that µn → µ weakly
and h(f (xn , un , w)) continuously converge to h(f (x, u, w)) i.e. h(f (xn , un , wn )) → h(f (x, u, w) for some wn → w
since f and g are continuous functions. Similarly, it can be also shown that Tn (·|x, u) and T (·|x, u) are weakly continuous
on (x, u). Thus, for the case where the channel is continuous in total variation by Theorem 12.2.8 and Theorem 12.2.4 if
c(x, u) is bounded and U is compact the result follows.
For the fully observed case, Jβ∗ (Tn ) → Jβ∗ (T ) by Theorem 12.3.2 and Jβ (T , γn∗ ) → Jβ∗ (T ) by Theorem 12.3.3. ⋄

Remark 12.14. We note here that the moment estimation method can also lead to consistency. Suppose that the distribution
of W is determined by its moments, such that estimate models Wn have moments of all orders and limn = E[Wnr ] =
E[W r ] for all r ∈ Z+ . Then, we have that [43, Thm 30.2] Wn → W weakly and thus Tn (·|xn , un ) → T (·|x, u) weakly
for any (xn , un ) → (x, u) under the assumptions of above corollary. Hence, we reach continuity and robustness using the
same arguments as in the previous result (Corollary 12.13).

Now, we give a similar result with the assumption that the noise process of the dynamics admits a continuous probability
density function.

Corollary 12.15 (to Theorem 12.2.6 and Theorem 12.2.7). Suppose we are given the following dynamics for real vector
state space X and action space U
xt+1 = f (xt , ut , wt ), yt = g(xt , vt ),
where {wt } and {vt } are [Link] processes and the noise models are unknown but it is known that the noise wt admits
a continuous probability density function. Suppose that f (x, u, ·) : W → X is invertible for all (x, u). We collect i.i.d.
samples of {wt } as in (12.14) and use them to construct an estimator, µ̃n , as described in [103] which consistently
estimates µ in total variation. Using these empirical estimates, we construct the empirical transition kernel estimates for
any (x0 , u0 ) and Borel B as

Tn (B|x0 , u0 ) = µ̃n (fx−1


0 ,u0
(B)).

Then independent of the channel, Jβ∗ (Tn ) → Jβ∗ (T ) and Jβ (T , γn∗ ) → Jβ∗ (T ), where γn∗ is the optimal policy designed
for Tn . Since the channel model g has no restrictions, this result also applies to the fully observed model setup by taking
g(xt , vt ) = xt .

Proof. By [103] we can estimate µ in total variation so that almost surely


limn→∞ ∥µ̃n − µ∥T V = 0. We claim that the convergence of µ̃n to µ under total variation metric implies the convergence
of Tn to T in total variation uniformly over all x ∈ X and u ∈ U i.e. limn→∞ supx,u ∥Tn (·|x, u) − T (·|x, u)∥T V = 0.
Observe the following:
12.5 Applications to Data-Driven Learning and Finite Model Approximations 313

sup ∥Tn (·|x, u) − T (·|x, u)∥T V


x,u
Z Z
= sup sup h(x1 )Tn (dx1 |x, u) − h(x1 )T (dx1 |x, u)
x,u ||h||∞ ≤1
Z Z
= sup sup h(f (x, u, w))µ̃n (dw) − h(f (x, u, w))µ(dw)
x,u ||h||∞ ≤1

≤∥µ̃n − µ∥T V → 0.

Thus, by Theorem 12.2.6 and Theorem 12.2.7, the result follows. ⋄

The following example presents some system and channel models which satisfy the requirements of the above corollaries.

Example 12.16. Let X, Y, U be real vector spaces with

xt+1 = f (xt , ut ) + wt , yt = h(xt , vt )

for unknown i.i.d. noise processes {wt } and {vt }.

1. Suppose the channel is in the following form; yt = h(xt , vt ) = xt + vt where vt admits a density (e.g. Gaussian
density). It can be shown by an application of Scheffé’s theorem that the channels in this form are continuous in
total variation. If further f (xt , ut ) is continuous and bounded then the requirements of Corollary 12.13 hold for
partially observed models.
2. If the channel is in the following form; xt = h(xt , vt ) then the system is fully observed. If further f (xt , ut ) is
continuous and bounded then the requirements of Corollary 12.13 holds for fully observed models.
3. Suppose the function f (xt , ut ) is known, if the noise process wt admits a continuous density, then one can estimate
the noise model in total variation in a consistent way (see [103]). Hence, the conditions of Corollary 12.15 holds
independent of the channel model.

12.5.2 Application to Approximations of MDPs and POMDPs with Weakly Continuous Kernels

We now discuss Problem P4, that is whether approximation of an MDP model with a standard Borel space with a finite
MDPs can be viewed an instance of robustness problem to incorrect models and whether our results can be applied.
In Section 8.2, we presented conditions under which finite state/action models are asymptotically optimal. Here, we view
those approximation results as an instance of robustness. We will focus on the weakly continuous model setup.
By Section 8.2.1 we know that finite quantization policies are nearly optimal under mild weak continuity conditions (see
Assumption 8.2.1). Thus, to make the presentation shorter, we will either assume that the action set is finite, or it has
been approximated by a finite action space through the construction above. Assuming finite action sets will help us avoid
measurability issues ( [275, p. 6-7]) as well as issues with existence of optimal policies.
One can write the following fixed point equation for the finite MDP
( )
X
n n
Jβ (x) = min cn (x, a) + β Jβ (x1 )Tn (x1 |x, a)
a∈U
x1 ∈Xn

where Tn is the transition model for the finite MDP and cn is the cost function defined on the finite model. Since the acton
space is finite, we can find an optimal policy, say fn∗ for this fixed point equation. One can also simply extend Jβn and fn∗ ,
which are defined on Xn to the entire state space X by taking them constant over the quantization bins Sn,i . If we call the
extended versions Jˆβn and fˆn , the following result holds, which is a re-statement of Theorem ??:
314 12 Robustness to Incorrect Models and Learning

Theorem 12.5.1 [275, Theorem 2.2 and 4.1] Suppose Assumption 8.2.1 holds. Then, for any β ∈ (0, 1) the discounted
cost of the deterministic stationary policy fˆn , obtained by extending the discounted optimal policy fn∗ of f-MDPm to X
(i.e., fˆn = fn∗ ◦ Qn ), converges to the discounted value function J ∗ of the compact-state MDP:

lim ∥Jˆβn (· ) − Jβ∗ (· )∥ = 0 and lim ∥Jβ (fˆn , · ) − Jβ∗ ∥ = 0. (12.15)


n→∞ n→∞

Theorem 12.5.1 shows that under Assumption 8.2.1, an optimal solution can be approximated via the solutions of finite
models. We now show that the above approximation scheme can be viewed in relation to our robustness results.
Proof sketch of Theorem 12.5.1 via results from Section 12.3. With the introduced setup, one can see that the extended
value function and optimal policy for the finite model satisfy the following:
 Z 
ˆn ˆn
Jβ (x) = min ĉn (x, u) + β Jβ (x1 )T̂n (dx1 |x, u)
a∈U

where ĉn is the extended version of cn to the state space X by making it constant over the quantization bins {Sn,i }i and T̂n
is such that for any function f
Z Z Z
f (x1 )T̂n (dx1 |x, u) := f (x1 )T (dx1 |z, u)ψn,i (dz)
x1 ∈X z∈Sn,i

where Sn,i is the quantization bin that x belongs to.


With this setup, one can see that for any xn → x we have ĉn (xn , u) → c(x, u) and for any continuous and bounded f
Z Z Z
f (x1 )T̂n (dx1 |xn , u) := f (x1 )T (dx1 |z, u)ψn,i (dz)
x1 ∈X z∈Sn,i
Z
→ f (x1 )T (dx1 |x, u).

Hence, Assumption 12.1.1 holds under Assumption 8.2.1, and we can conclude the proof using Theorem 12.3.3 and
Theorem 12.3.2. ⋄

12.6 Bibliographic Notes

In this chapter, we studied regularity properties of optimal stochastic control on the space of transition kernels, and ap-
plications to robustness of optimal control policies designed for an incorrect model applied to an actual system. We also
presented applications to data-driven learning and related the robustness problem to finite MDP approximation techniques.
For the problems presented in this chapter, our focus was on infinite horizon discounted cost setup. However, we note that
the results can be extended to the infinite horizon average cost setup under various forms of ergodicity properties on the
state process.
Robustness is a desired property for the optimal control of stochastic or deterministic systems when a given model does
not reflect the actual system perfectly, as is usually the case in practice. This is a classical problem, and there is a very
large literature on robust stochastic control and its application to learning-theoretic methods; see e.g. [24, 31, 52, 117,
147, 169, 337], [6, 18, 114, 118, 178, 185, 186, 210, 240, 250, 254, 298, 311]. Studies on robustness via minimax methods
include [176, 245]. A comprehensive literature review is presented in [182, 186]. For empirical learning methods and their
stability properties, see [112, 148]
This chapter primarily builds on [181, 182, 186, 191].
12.7 Exercises 315

12.7 Exercises

Exercise 12.7.1
A

Basics of Function Spaces

A.1 Normed Linear (Vector) Spaces and Metric Spaces

Definition A.1.1 A linear (vector) space X is a space which is closed under addition and scalar multiplication: In partic-
ular, we define an addition operation, + and a scalar multiplication operation · such that

+:X×X→X

·:C×X→X
with the following properties (we note that we may take the scalars to be either real or complex numbers). The following
are satisfied for x, y ∈ X and α, β scalars:
(i) x + y = y + x
(ii) (x + y) + z = x + (y + z).
(iii) α · (x + y) = α · x + α · y.
(iv) (α + β) · x = α · x + β · x.
(v) There is a null vector 0 such that x + 0 = x.
(vi) α · (β · x) = (αβ) · x
(vii) For every x ∈ X, 1 · x = x
(viii) For every x ∈ X, there exists an element, called the (additive) inverse of x and denoted with −x with the property
x + (−x) = 0.

Example A.1. (i) The space Rn is a linear space. The null vector is 0 = (0, 0, · · · , 0) ∈ Rn .
(ii) Consider the interval [a, b]. The collection of real-valued continuous functions on [a, b] is a linear space. The null
element 0 is the function which is identically 0. This space is called the space of real-valued continuous functions on [a, b]
(iii) The set of all infinite sequences of real numbers having only a finite number of terms not equal to zero is a vector
space. If one adds two such sequences, the sum also belongs to this space. This space is called the space of finitely many
non-zero sequences.
(iv) The collection of all polynomial functions defined on an interval [a, b] with complex coefficients forms a complex
linear space. Note that the sum of polynomials is another polynomial.

Definition A.1.2 A non-empty subset M of a (real) linear vector space X is called a subspace of X if

αx + βy ∈ M, ∀x, y ∈ M and α, β ∈ R.
318 A Basics of Function Spaces

In particular, the null element 0 is an element of every subspace. For M, N two subspaces of a vector space X, M ∩ N is
also a subspace of X.

Definition A.1.3 A normed linear space X is a linear vector space on which a map from X to R, that is a member of
Γ (X; R)) called norm is defined such that:
– ||x|| ≥ 0 ∀x ∈ X, ||x|| = 0 if and only if x is the null element (under addition and multiplication) of X.
– ||x + y|| ≤ ||x|| + ||y||
– ||αx|| = |α|||x||, ∀α ∈ R, ∀x ∈ X

Definition A.1.4 In a normed linear space X, an infinite sequence of elements {xn } converges to an element x if the
sequence {||xn − x||} converges to zero.

Example A.2. a) The normed linear space C([a, b]) consists of continuous functions on [a, b] together with the norm ||x|| =
max{a≤t≤b} |x(t)|.
  p1
p
P
b) lp (Z+ ; R) := {x ∈ Γ (Z+ ; R) : ||x||p = i∈Z+ |x(i)| < ∞} is a normed linear space for all 1 ≤ p < ∞. c)
Recall that if S is a set of real numbers bounded above, then there is a smallest real number y such that x ≤ y for all x ∈ S.
The number y is called the least upper bound or supremum of S. If S is not bounded from above, then the supremum is ∞.
In view of this, for p = ∞, we define

l∞ (Z+ ; R) := {x ∈ Γ (Z+ ; R) : ||x||∞ = sup |x(i)| < ∞}


i∈Z+

  p1
Rb
d) Lp ([a, b]; R) = {{x ∈ Γ ([a, b]; R) : ||x||p = a
|x(i)|p < ∞} is a normed linear space. For p = ∞, we typically
write: L∞ ([a, b]; R) := {x ∈ Γ ([a, b]; R) : ||x||∞ = supt∈[a,b] |x(t)| < ∞}. However, for 1 ≤ p < ∞, to satisfy the
condition that ||x||p = 0 implies that x(t) = 0, we need to assume that functions which are equal to zero almost everywhere
are equivalent; for p = ∞ the definition is often revised with essential supremum instead of supremum so that

||x||∞ = inf sup |y(t)|


y(t)=x(t)a.e. t∈[a,b]

To show that lp defined above is a normed linear space, we need to show that ||x + y||p ≤ ||x||p + ||y||p .

Theorem A.1.1 (Minkowski’s Inequality) For 1 ≤ p ≤ ∞

||x + y||p ≤ ||x||p + ||y||p

The proof of this result uses a very important inequality, known as Hölder’s inequality.

Theorem A.1.2 (Hölder’s Inequality) X


x(k)y(k) ≤ ||x||p ||y||q ,
with 1/p + 1/q = 1 and 1 ≤ p, q ≤ ∞.

Definition A.1.5 A metric defined on a set X, is a function d : X × X → R such that:


– d(x, y) ≥ 0, ∀x, y ∈ X and d(x, y) = 0 if and only if x = y.
– d(x, y) = d(y, x), ∀x, y ∈ X.
– d(x, y) ≤ d(x, z) + d(z, y), ∀x, y, z ∈ X.
A.1 Normed Linear (Vector) Spaces and Metric Spaces 319

Definition A.1.6 A metric space (X, d) is a set equipped with a metric d.

A normed linear space is also a metric space, with metric

d(x, y) = ||x − y||.

An important class of normed spaces that are widely used in optimization and engineering problems are Banach spaces:

A.1.1 Banach Spaces

Definition A.1.7 A sequence {xn } in a normed space X is Cauchy if for every ϵ, there exists an N such that ||xn −xm || ≤ ϵ,
for all n, m ≥ N .

The important observation on Cauchy sequences is that, every converging sequence is Cauchy, however, not all Cauchy
sequences are convergent: This is because the limit might not live in the original space where the sequence elements take
values in. This brings the issue of completeness:

Definition A.1.8 A normed linear space X is complete, if every Cauchy sequence in X has a limit in X. A complete normed
linear space is called Banach.

Banach spaces are important for many reasons including the following one: In optimization problems, sometimes we
would like to see if a sequence converges, for example if a solution to a minimization problem exists, without knowing
what the limit of the sequence could be. Banach spaces allow us to use Cauchy sequence arguments to claim the existence
of optimal solutions. If time allows, we will discuss how this is used by using contraction and fixed point arguments for
transformations.
In applications, we will also discuss completeness of a subset. A subset of a Banach space X is complete if and only if it is
closed. If it is not closed, one can provide a counterexample sequence which does not converge. If the set is closed, every
Cauchy sequence in this set has a limit in X and this limit should be a member of this set, hence the set is complete.

Exercise A.1.1 The space of bounded functions {x : [0, 1] → R, supt∈[0,1] |x(t)| < ∞} is a Banach space.

The above space is also denoted by L∞ ([0, 1]; R) or L∞ ([0, 1]).

  p1
p
P
Theorem A.1.3 lp (Z+ ; R) := {x ∈ f (Z+ ; R) : ||x||p = i∈N+ |x(i)| < ∞} is a Banach space for all 1 ≤ p ≤ ∞.

Sketch of Proof: The proof is completed in three steps.


(i) Let {xn } be Cauchy. This implies that for every ϵ > 0, ∃N such that for all n, m ≥ N ||xn − xm || ≤ ϵ. This also
implies that for all n > N , ||xn || ≤ ||xN || + ϵ. Now let us denote xn by the vector {xn1 , xn2 , xn3 . . . , }. It follows that for
every k the sequence {xnk } is also Cauchy. Since xnk ∈ R, and R is complete, xnk → xk for some xk . Thus, the sequence
xn pointwise converges to some vector x∗ .

(ii) Is x ∈ lp (Z+ ; R)? Define xn,K = {xn1 , xn2 , . . . , xnK−1 , xnK , 0, 0, . . . }, that is vector which truncates after the Kth
coordinate. Now, it follows that
||xn,K || ≤ ||xN || + ϵ,
for every n ≥ N and K and
K
X K
X
p
lim ||xn,K || = lim |xni |p = |xi |p ,
n→∞ n→∞
i=1 i=1
320 A Basics of Function Spaces

since there are only finitely many elements in the summation. The question now is whether ||x∞ || ∈ p(Z+ ; R). Now,

||xn,K || ≤ ||xN || + ϵ,

and thus
lim ||xn,K || = ||xK || ≤ ||xN || + ϵ,
n→∞

Let us take another limit, by the monotone convergence theorem (Recall that this theorem says that a monotonically
increasing sequence which is bounded has a limit).
K
X
lim ||x∗,K ||p = lim |xi |p = ||x∞ ||pp ≤ ||xN || + ϵ.
K→∞ K→∞
i=1

(iii) The final question is: Does ||xn − x∗ || → 0? Since the sequence is Cauchy, it follows that for n, m ≥ N

||xn − xm || ≤ ϵ

Thus,
||xn,K − xm,K || ≤ ϵ
and since K is finite
lim ||xn,K − xm,K || = ||xn,K − x∗,K || ≤ ϵ
m→∞

Now, we take another limit


lim ||xn,K − x∗,K || ≤ ϵ
K→∞

By the monotone convergence theorem again,

lim ||xn,K − x∗,K || = ||xn − x|| ≤ ϵ


K→∞

Hence, ||xn − x|| → 0. ⋄


The above spaces are also denoted lp (Z+ ), when the range space is clear from context.
The following is a useful result.

Theorem A.1.4 (Hölder’s Inequality) X


x(t)y(t) ≤ ||x||p ||y||q ,
with 1/p + 1/q = 1 and 1 ≤ p, q ≤ ∞.

Remark: A brief remark for notations: When the range space is R, the notation lp (Ω) denotes lp (Ω; R) for a discrete-time
index set Ω and likewise for a continuous-time index set Ω, Lp (Ω) denotes Lp (Ω; R). ⋄

A.1.2 Hilbert Spaces

We first define pre-Hilbert spaces.

Definition A.1.9 A pre-Hilbert space X is a linear vector space where an inner product is defined on X ×X. Corresponding
to each pair x, y ∈ X the inner product ⟨x, y⟩ is a scalar (that is real-valued or complex-valued). The inner product satisfies
the following axioms:
1. ⟨x, y⟩ = ⟨y, x⟩∗ (the superscript denotes the complex conjugate) (we will also use ⟨y, x⟩ to denote the complex
conjugate)
2. ⟨x + y, z⟩ = ⟨x, z⟩ + ⟨y, z⟩
A.1 Normed Linear (Vector) Spaces and Metric Spaces 321

3. ⟨αx, y⟩ = α⟨x, y⟩
4. ⟨x, x⟩ ≥ 0, equals 0 iff x is the null element.

The following is a crucial result in such a space, known as the Cauchy-Schwarz inequality, the proof of which was presented
in class:

Theorem A.1.5 For x, y ∈ X, p p


⟨x, y⟩ ≤ ⟨x, x⟩ ⟨y, y⟩,
where equality occurs if and only if x = αy for some scalar α.
p
Exercise A.1.2 In a pre-Hilbert space ⟨x, x⟩ defines a norm: ||x|| = ⟨x, x⟩

p
The proof for the result requires one to show that ⟨x, x⟩ satisfies the triangle inequality, that is

||x + y|| ≤ ||x|| + ||y||,

which can be proven by an application of the Cauchy-Schwarz inequality.


P
Not all spaces admit an inner product. In particular, however, l2 (N+ ; R) admits an inner product with ⟨x, y⟩ = t∈N+ x(n)y(n)
p
for x, y ∈ l2 (N+ ; R). Furthermore, ||x|| = ⟨x, x⟩ defines a norm in l2 (N+ ; R).
The inner product, in the special case of RN , is the usual inner vector product; hence RN is a pre-Hilbert space with the
usual inner-product.

Definition A.1.10 A complete pre-Hilbert space, is called a Hilbert space.

Hence, a Hilbert space is a Banach space, endowed with an inner product, which induces its norm.

Proposition A.1.1 The inner product is continuous: if xn → x, and yn → y, then ⟨xn , yn ⟩ → ⟨x, y⟩ for xn , yn in a Hilbert
space.

Proposition A.1.2 In a Hilbert space X, two vectors x, y ∈ X are orthogonal if ⟨x, y⟩ = 0. A vector x is orthogonal to a
set S ⊂ X if ⟨x, y⟩ = 0 ∀y ∈ S.

Theorem A.1.6 (Projection Theorem:) Let H be a Hilbert space and B a closed subspace of H. For any vector x ∈ H,
there is a unique vector m ∈ B such that

||x − m|| ≤ ||x − y||, ∀y ∈ B.

A necessary and sufficient condition for m ∈ B to be the minimizing element in B is that, x − m is orthogonal to B.

A.1.3 Separability

Definition A.1.11 Given a normed linear space X, a subset D ⊂ X is dense in X, if for every x ∈ X, and each ϵ > 0,
there exists a member d ∈ D such that ||x − d|| ≤ ϵ.

Definition A.1.12 A set is countable if every element of the set can be associated with an integer via an ordered mapping.

Examples of countables spaces are finite sets and the set Q of rational numbers. An example of uncountable sets is the set
R of real numbers.
322 A Basics of Function Spaces

Theorem A.1.7 a) A countable union of countable sets is countable. b) Finite Cartesian products of countable sets is
countable. c) Infinite Cartesian products of countable sets may not be countable. d) [0, 1] is not countable.

Cantor’s diagonal argument and the triangular enumeration are important steps in proving the theorem above.
Since rational numbers are the ratios of two integers, one may view rational numbers as a subset of the product space of
countable spaces; thus, rational numbers are countable.

Definition A.1.13 A space X is separable, if it contains a countable dense set.

Separability informs us that for approximation purposes it suffices to work with a countable set, when the set is uncountable.
Examples of separable sets are R, and the set of continuous and bounded functions on a compact set metrized with the
maximum distance between the functions.
Complete, separable and metrizable spaces form a very broad class of signal spaces. Such spaces are called Polish metric
spaces when a metric is defined apriori. Borel subsets of such spaces are called standard Borel spaces.
B

On the Convergence of Random Variables

B.1 Limit Events and Continuity of Probability Measures

Given A1 , A2 , . . . , An , · · · ∈ F, define:
lim sup An = ∩∞ ∞
n=1 ∪k=n Ak
n

lim inf An = ∪∞ ∞
n=1 ∩k=n Ak
n

For the superior limit, an element is in this set, if it is in infinitely many An s. For the inferior case, an element is in the
limit, if it is in almost except for a finite number of An s. The limit of a sequence of sets exists if the above limits are equal.
We have the following result:

Theorem B.1.1 For a sequence of events An :

P (lim inf An ) ≤ lim inf P (An ) ≤ lim sup P (An ) ≤ P (lim sup An )
n n n n

We have the following regarding continuity of probability measures:

Theorem B.1.2 (i)For a sequence of events An with An ⊂ An+1 for all n,

lim P (An ) = P (∪∞


n=1 An )
n→∞

(ii) For a sequence of events An with An+1 ⊂ An for all n,

lim P (An ) = P (∩∞


n=1 An )
n→∞

B.2 Borel-Cantelli Lemma


P∞ P
Theorem B.2.1 (i) If n=1 P (An ) < ∞, then P (lim supn An ) = 0. (ii) If {An } are independent and if P (An ) = ∞,
then P (lim supn An ) = 1.

P∞
Proof sketch. (i) For every M ∈ N: P (lim supn An ) = P (∩∞ ∞ ∞
n=1 ∪k=n Ak ) ≤ P (∪k=M Ak ) ≤ k=M P (Ak ). Therefore,
as M → ∞, the sum on the right is less than any Q∞ ϵ > 0. Since this
Q∞ϵ is arbitrary, the result follows. (ii) Note that for every
M ∈ N, P ((∪∞ k=M A k )c
) = P (∩∞
k=M kA c
) = k=M P (A c
k ) = k=M (1 − P (A k )), where we use independence of the
events. Thus, by Exercise 3.5.7, we have that P ((∪∞ k=M A k ) c
) = 0. Since for every M , P (∪ ∞
k=M A k ) = 1, and the sets
∪∞k=M A k are non-expanding, the result follows from Theorem B.1.2. ⋄
324 B On the Convergence of Random Variables

Exercise B.2.1 Let {An } be a sequence of independent events where An is the event that the nth coin flip is head. What is
the probability that there are infinitely many heads if P (An ) = 1/n2 ?

An important application of the above is the following:

Theorem B.2.2 Let Zn , n ∈ N and Z be random variables and for every ϵ > 0,
X
P (|Zn − Z| ≥ ϵ) < ∞.
n

Then,
P ({ω : Zn (ω) = Z(ω)}) = 1.
That is Zn converges to Z with probability 1.

B.3 Convergence of Random Variables

B.3.1 Convergence almost surely (with probability 1)

Definition B.3.1 A sequence of random variables Xn converges almost surely to a random variable X if P ({ω :
limn→∞ Xn (ω) = X(ω)}) = 1.

B.3.2 Convergence in Probability

Definition B.3.2 A sequence of random variables Xn converges in probability to a random variable X if limn→∞ P (|Xn −
X| ≥ ϵ}) = 0 for every ϵ > 0.

B.3.3 Convergence in Mean-square

Definition B.3.3 A sequence of random variables Xn converges in the mean-square sense to a random variable X if
limn→∞ E[|Xn − X|2 ] = 0.

B.3.4 Convergence in Distribution

Definition B.3.4 Let Xn be a random variable with cumulative distribution function Fn , and X be a random variable with
cumulative distribution function F . A sequence of random variables Xn converges in distribution (or weakly) to a random
variable X if limn→∞ Fn (x) = F (x) for all points of continuity of F .

Theorem B.3.1 a) Convergence in almost sure sense implies in probability. b) Convergence in mean-square sense implies
convergence in probability. c) If Xn → X in probability, then Xn → X in distribution.

We also have partial converses for the above results:

Theorem B.3.2 a) If P (|Xn | ≤ Y ) = 1 for some random variable Y with E[Y 2 ] < ∞, and if Xn → X in probability,
then Xn → X in mean-square. b) If Xn → X in probability, there exists a subsequence Xnk which converges to X almost
surely. c) If Xn → X and Xn → Y in probability, mean-square, or almost surely; then P (X = Y ) = 1.
B.3 Convergence of Random Variables 325

A sequence of random variables is uniformly integrable if:

lim sup E[|Xn |1|Xn |≥K ] = 0.


K→∞ n

A sufficient condition for a sequence of random variables to be uniformly integrable is that there exists a function g : R →
|Xn |g(|Xn |)
R with the property as t → ∞, g(t) t ↑ ∞, so that supn E[g(Xn )] < ∞: (to see this, note that g(|Xn |) = |Xn | ≥
g(K)
K |Xn |, for |Xn | ≥ K. Thus,

E[g(Xn )1|Xn |≥K ] E[g(Xn )]


lim sup E[|Xn |1|Xn |≥K ] ≤ lim sup ≤ lim sup → 0.
K→∞ n K→∞ n g(K)/K K→∞ n g(K)/K

Note also that, if {Xn } is uniformly integrable, then, supn E[|Xn |] < ∞.

Theorem B.3.3 Under uniform integrability, convergence in almost sure sense implies convergence in mean-square.

Theorem B.3.4 If Xn → X in probability, there exists some subsequence Xnk which converges to X almost surely.

A further useful result is the following.

Theorem B.3.5 [Skorohod’s representation theorem] Let Xn → X in distribution. Then, there exists a sequence of random
variables Yn and Y such that, Xn and Yn have the same cumulative distribution functions; X and Y have the same
cumulative distribution functions and Yn → Y almost surely.

With the above, we can prove the following result.

Theorem B.3.6 The following are equivalent: i) Xn converges to X in distribution. ii) E[f (Xn )] → E[f (X)] for all
continuous and bounded functions f . iii) The characteristic functions Φn (u) := E[eiuXn ] converge pointwise for every
u ∈ R.
C

Some Remarks on Measurable Selections

As we observe in Chapter 5, in stochastic control measurability issues arise extensively both for the measurability of control
policies as well as that of value functions/optimal costs. Theorem 5.1.1 and 5.2.1, and Lemma 5.2.4 are some examples
where these were crucially utilized. In addition, we observed that the theory of martingales and filtration, the measurability
properties are essential.
One particular aspect is to ensure that maps of the form:

J(x) := inf c(x, u) (C.1)


u∈U

are measurable or at least Lebesgue-integrable.

Theorem C.0.1 [Kuratowski Ryll-Nardzewski Measurable Selection Theorem] [201] [283] and [168, Theorem 2] Let X, U
be Polish spaces and Γ = {(x, ψ(x)), x ∈ X} where ψ(x) ⊂ U be such that, ψ(x) is closed for each x ∈ X and Γ be a
Borel measurable set in X × U. Then, there exists at least one measurable function f : X → U such that {(x, f (x)), x ∈
X} ⊂ Γ .

A proof sketch is as follows for the case with U = R+ and ψ(x) is compact valued. With n ∈ N, consider the infinite
sequence of rationals {k/n; k ∈ Z+ }. Consider ψ −1 ([ nk , k+1
n )). Define the Borel set

k k+1 −1 k k + 1
X(k,n) = ψ −1 ([ , )) \ ∪k−1
m=1 ψ ([ , )).
n n n n
Then, define a multi-function:
k k+1
ψ n (x) = ψ(x) ∩ [ , ),
n n
whenever x ∈ X(k,n) . Now, each ψ n is a multi-function. Take n → ∞, in this case, since ψ(x) is closed, each converging
subsequence unk ∈ X(k,n) is so that limnk →∞ unk ∈ ψ(x) (by the closed property). Therefore, for each x, the limit
limn→∞ ψ n (x) is well-defined and single-valued, and is placed in ψ(x). This approach can be generalized.
This result was utilized in Chapter 5 (see Lemma 5.2.4). Recall also the relationship of the argument with that in the proof
of Theorem 5.1.1.
In the following, we assume that the spaces considered are Polish. A function f is µ-measurable if there exists a Borel
measurable function g which agrees with f µ-a.e. A function that is µ-measurable for every probability measure is called
universally measurable.
A measurable image of a Borel set is called an analytic set [116].

Fact C.0.1 The image of a Borel set under a measurable function, and hence an analytic set, is universally measurable.
328 C Some Remarks on Measurable Selections

Remark C.1. We note that in some texts, an analytic set is defined as the continuous image of a Borel set. However, as [116]
notes, one could always express the image of a Borel set A under a measurable function f : X → Y as a projection (which
is a continuous map) of (A, f (A)) onto Y.

The integral of a universally measurable function is well-defined and is equal to the integral of a Borel measurable function
which is µ-almost equal to that function. While applying dynamic programming, we often seek to establish the existence
of measurable functions through the operation:
 Z 
Jt (xt ) = inf c(x, u) + Jt+1 (xt+1 )Q(dxt+1 |xt , u)
u∈U(xt )

However, we need a stronger condition that universal measurability for the recursions to be well-defined. A function f is
called lower semi-analytic if {x : f (x) < c} is analytic for each scalar c.

Theorem C.0.2 [116] Let i : X → 2S (that is, i maps X to subsets of S) be such that i−1 is Borel measurable, and
f : S → R be measurable. Then:
v(x) = inf f (z)
z:z∈i(x)

is lower semi-analytic.

Observe that (see p. 85 of [116])


{x : v(x) < c} = i−1 ({z : f (z) < c})
The set {z : f (z) < c} is Borel, and thus if i−1 is also Borel, it follows that v is lower semi-analytic. We require then that
i−1 : S → X to be Borel. Consider now the following application.

Theorem C.0.3 Consider G = {(x, u) : u ∈ U(x)} which is a Borel measurable set. The map,

v(x) = inf v(x, z),


(x,z)∈G

is lower semi-analytic.

Proof. The graph G is measurable. It follows that{x : v(x) < c} = i−1 ({(x, z) : v(x, z) < c}), where i−1 is the projection
of G onto X, which is a continuous operation; the image may not be measurable but as a measurable mapping of a Borel
set, it is analytic. As a result v is lower semi-analytic. ⋄

Theorem C.0.4 Lower semi-analytic functions are universally measurable.

Implication: Dynamic programming can be carried out for such expressions. In particular, the following is due to Bertsekas
and Shreve [37, Chapter 7]:

Theorem C.0.5 The following hold:


(i) Let E1 , E2 be Borel and g : E1 × E2 → R be lower semi-analytic. Then,

h(e1 ) = inf g(e1 , e2 )


e2 ∈E2

is lower semi-analytic.
(ii) Let E1 , E2 be Borel and g : E1 × E2 → R be lower semi-analytic. Let Q(de2 |e1 ) be a stochastic kernel. Then,
Z
f (e1 ) := g(e2 )Q(de2 |e1 )

is lower semi-analytic.
C Some Remarks on Measurable Selections 329

We note that the second result would not be correct if g is only taken to be universally measurable. The result above ensures
that we can follow the dynamic programming arguments in an inductive manner under conditions that are less restrictive
than the conditions stated in the measurable selection conditions. These then imply the existence of ϵ-optimal solutions
(possibly universally measurable) [288].
Building on this discussion, and the material in Chapter 5, we summarize three useful results in the following.

Fact C.0.2 Consider (C.1).


(i) If c is continuous on X × U and U is compact, then J is continuous and there exists an optimal measurable policy.
(ii) If the measurable c is continuous on U for every x, and U is compact, then J is measurable (Prop D.5 in [164] and
Himmelberg and Schäl [283]); see Theorem 5.2.4. Furthermore, there exists an optimal measurable policy.
(iii) [37, Prop. 7.47 and 7.50] If c is measurable on X × U and U is Borel, then J is lower semi-analytic. Furthermore,
there exists a near optimal universally measurable function.

Compactness of U is a crucial component for some of these results. However, as discussed in Section 5.2, U(x) may be
allowed to depend on x, for item (ii) under the assumption that the graph G = {(x, u) : u ∈ U(x)} defined above is
Borel, and U(x) is compact for every x; see p. 182 in [164] (and also [168], [283], [124] and [201], among others). For (i),
this relaxation also requires that the set valued map U(x) is upper semi-continuous: let xn → x, then for every sequence
un ∈ U(xn ), there exists a subsequence which converges to some u where every such limit u satisfies u ∈ U(x).
D

On Spaces of Probability Measures

In this section we present various topologies, and when applicable, several metrics on the sets of probability measures.

D.1 Convergence of Sequences of Probability Measures

Let X be a Polish space and let P(X) denote the family of all probability measures on (X, B(X)). Let {µn , n ∈ N} be a
sequence in P(X).
The sequence {µn } is said to converge to µ ∈ P(X) weakly if
Z Z
c(x)µn (dx) → c(x)µ(dx) (D.1)
X X

for every continuous and bounded c : X → R. 1


On the other hand, {µn } is said to converge to µ ∈ P(X) setwise if
Z Z
c(x)µn (dx) → c(x)µ(dx)
X X

for every measurable and bounded c : X → R. Setwise convergence can also be defined through pointwise convergence on
Borel subsets of X (see, e.g., [167]), that is

µn (A) → µ(A), for all A ∈ B(X)

since the space of simple functions are dense in the space of bounded and measurable functions under the supremum norm.
For two probability measures µ, ν ∈ P(X), the total variation metric is given by

∥µ − ν∥T V := 2 sup |µ(B) − ν(B)|


B∈B(X)
Z Z
= sup f (x)µ(dx) − f (x)ν(dx) , (D.2)
f : ∥f ∥∞ ≤1

where the supremum is over all measurable real f such that ∥f ∥∞ = supx∈X |f (x)| ≤ 1. A sequence {µn } is said to
converge to µ ∈ P(X) in total variation if ∥µn − µ∥T V → 0.
1
It is important to emphasize that what is typically studied in probability as weak convergence is not the exact weak convergence
notion used in functional analysis: The topological dual space of the set of probability measures does not only consist of expectations of
continuous and bounded functions. However, the dual space of the space of continuous and bounded functions with the supremum norm
does admit a representation in terms of expectations [219]; hence, the weak convergence here is in actuality the weak∗ convergence in
analysis and distribution theory.
332 D On Spaces of Probability Measures

Setwise convergence is equivalent to pointwise convergence on Borel sets whereas total variation requires uniform conver-
gence on Borel sets. Thus these three convergence notions are in increasing order of strength: convergence in total variation
implies setwise convergence, which in turn implies weak convergence.
On the other hand, total variation is a stringent notion for convergence. For example a sequence of discrete probability
measures never converges in total variation to a probability measure which admits a density function with respect to
the Lebesgue measure and such a space is not separable. Setwise convergence also induces a topology on the space of
probability measures and channels which is not easy to work with since the space under this convergence is not metrizable
[142, p. 59].
However, the space of probability measures on a complete, separable, metric (Polish) space endowed with the topology of
weak convergence is itself a complete, separable, metric space [42].
There are various ways to metrize weak convergence. One immediate metric builds on the following reasoning: One can
construct (since the space of continuous functions on a compact set is separable under the supremum norm) a countable
collection of continuous functions {ck , k ∈ N} such that it suffices to only consider these functions in (D.1) to establish
weak-convergence. We can thus use these weak-convergence R determiningR functions (see e.g. [119, Theorem 3.4.5]) to
define a countable collection of semi-norms dk (µ, ν) := | ck (x)µ(dx) − ck (x)ν(dx|, and from these we can construct
a locally convex space which is metrizable. Thus, we have the following metric which metrizes the weak topology:

X Z Z
−(m+1)
ρ(µ, ν) = 2 fm (x)µ(dx) − fm (x)ν(dx) , (D.3)
m=1 S S

where {fm }m≥1 is an appropriate sequence of continuous and bounded functions such that ∥fm ∥∞ ≤ 1 for all m ≥ 1
(see [249, Theorem 6.6, p. 47]).
The Prohorov metric [42] also can be used to metrize this convergence topology.
As a more practical metric, the Wasserstein metric can also be used (for compact X) to metrize the weak convergence space
topology.

Definition D.1.1 (Wasserstein metric) The Wasserstein metric of order p, 1 ≤ p < ∞, for two distributions µ, ν ∈ P(X)
with finite pth moments (thus defined only on such a subset of P(X)) is defined as
Z  p1
Wp (µ, ν) = inf η(dx, dy)∥x − y∥p ,
η∈H(µ,ν) X×X

where H(µ, ν) denotes the set of probability measures on X × X with first marginal µ and second marginal ν, and ∥ · · · ∥
is a norm.

For compact X, the Wasserstein distance of order p metrizes the weak topology on the set of probability measures on X (see
[320, Theorem 6.9]; one can also see the connection viaR Theorem B.3.5).RFor non-compact X, weak convergence combined
with convergence of moments up to order p (that is of µn (dx)∥x∥p → µ(dx)∥x∥p ) is equivalent to convergence in Wp .
Finally, the bounded-Lipschitz metric ρBL [320, p.109] can also be used to metrize weak convergence:
Z Z
ρBL (µ, ν) = sup f (e)µ(de) − f (e)ν(de) , (D.4)
∥f ∥BL ≤1 X X

where
f (e) − f (e′ )
∥f ∥BL := ∥f ∥∞ + sup ,
e̸=e′ dX (e, e′ )

and dX is the metric on X.


We note that W1 can equivalently be written as [320, Remark 6.5]:
D.3 A Generalized Dominated Convergence Theorem 333
Z Z
W1 (µ, ν) := sup f (e)µ(de) − f (e)ν(de) ,
∥f ∥Lip ≤1 X X

where
f (e) − f (e′ )
∥f ∥Lip := sup .
e̸=e′ dX (e, e′ )

Comparing this with (D.4), it follows that

ρBL ≤ W1 . (D.5)

Another important distance measure (though not a metric) that is commonly used is relative entropy:

dP
R
Definition D.1.2 For two probability measures P and Q, relative entropy is defined as D(P ∥Q) = log dQ dP =
R dP dP dP
dQ log dQ dQ where P ≪ Q and dQ denotes the Radon-Nikodym derivative of P with respect to Q.

q
2
Total variation is related to relative entropy via Pinsker’s inequality [93]: ∥P −Q∥T V ≤ log(e) D(P ∥Q). This also shows
that convergence in relative entropy implies that under total variation.
Weak convergence is very important in applications of stochastic control and probability in general. Prohorov’s theorem
[110] provides a way to characterize compactness properties under weak convergence.

D.2 Some Measurability Results on Spaces of Probability Measures

Weak convergence topology leads to important measurability properties, as we discuss in the following two theorems. The
first one appears in [4] (see Theorem 15.13 in [4] or p. 215 in [53]).

Theorem D.2.1 Let S be a Polish space and M be the set of all measurable and bounded functions f : S → R. Then, for
any f ∈ M , the integral Z
π(dx)f (x)

defines a measurable function on P(S) under the topology of weak convergence.

This is a useful result since it allows us to define measurable functions in integral forms on the space of probability measures
when we work with the topology of weak convergence. The second useful result follows from Theorem D.2.1, [109,
Theorem 2.1] and [37, Proposition 7.25].

Theorem D.2.2 Let S be a Polish space. A function F : P(S) → P(S) is measurable on B(P(S)) (under weak conver-
gence), if for all B ∈ B(S) (F (·))(B) : P(S) → R is measurable under weak convergence on P(S), that is for every
B ∈ B(S), (F (π))(B) is a measurable function when viewed as a function from P(S) to R.

D.3 A Generalized Dominated Convergence Theorem

Under weak and setwise convergences, we can arrive at generalized forms of the dominated convergence theorem. In
particular, from [211, Theorem 3.5] and [287, Theorem 3.5], we have the following:

Theorem D.3.1 The following hold:


334 D On Spaces of Probability Measures

(i) Suppose that {µn }n ⊂ P(X) converges weakly to some µ. For a bounded real valued sequence of functions {fn }n
such that ∥fn ∥∞ < C for all n > 0 with C < ∞, if limn→∞ fn (xn ) = f (x) for all xn → x, i.e. fn continuously
converges to f , then Z Z
lim fn (x)µn (dx) = f (x)µ(dx).
n→∞ X X

(ii) Suppose that {µn }n ⊂ P(X) converges setwise to some µ. For a bounded real valued sequence of functions {fn }n
such that ∥fn ∥∞ < C for all n > 0 with C < ∞, if limn→∞ fn (x) = f (x) for all x, i.e. fn pointwise converges
to f , then Z Z
lim fn (x)µn (dx) = f (x)µ(dx).
n→∞ X X

D.4 The w-s Topology

Let, as before, X and Y be Polish spaces.

Definition
R D.4.1 The w-s topology on the set of probability measures P(X × Y) is the coarsest topology under which
f (x, y)µ(dx, dy) : P(X × Y) → R is continuous for every measurable and bounded f (x, y) which is continuous in y
for every x (but unlike the weak topology, f does not need to be continuous in x).

Theorem D.4.1 [284, Theorem 3.10] [27, Theorem 2.5] Let µn ∈ P(X × Y). If µn → µ weakly where the marginals
µn (dx × Y) →µ(dx × Y) setwise, then the convergence µn → µ is also in the w-s sense.

D.5 Lusin’s Theorem

Lusin’s theorem is a very consequential result in mathematical analysis.

Theorem D.5.1 [110, Theorem 7.5.2] Let (X, T ) be any topological space and µ a finite, closed regular Borel measure
on X. Let (S, d) be a separable metric space and let f be a Borel-measurable function from X into S. Then for any ϵ > 0
there is a closed set F ⊂ X such that µ(X \ F ) < ϵ and the restriction of f to F is continuous.

We also recall Tietze’s extension theorem, which is often used in conjunction with Lusin’s theorem to construct a continu-
ous extension of the continuous function defined on F in Theorem D.5.1 to X.

Theorem D.5.2 [113, Theorem 4.1][Tietze’s extension theorem] Let X be an arbitrary metric space, A a closed subset of
X, L a locally convex linear space, and f : A → L a continuous map. Then there exists a continuous function fC : X → L
such that fC (a) = f (a) ∀a ∈ A. Furthermore, the image of fC satisfies fC (X) ⊂ [convex hull of f (A)].
E

Relaxed Control Topologies

In deterministic as well as stochastic control theory, relaxed or randomized control policies allow for versatility in mathe-
matical analysis, leading to continuity, compactness, convexity and approximation properties, in a variety of system models,
cost criteria, and information structures.
Within the relaxed/randomized control framework, with X a state space, U a control space and with an X-valued random
variable X ∼ µ, instead of considering the set of deterministic admissible policies:
 
Γ = γ : γ is a measurable function from X to U , (E.1)

one considers
 
ΓR = γ : γ is a measurable function from X to P(U) , (E.2)

where P(U) is endowed with the Borel σ-algebra generated by the weak convergence topology.
On ΓR , two commonly studied topologies are the following.

E.1 Young Topology on Control Policies

A prominent approach since Young’s seminal paper [340] has been via the study of topologies on Young measures defined
by randomized/relaxed controls, where one views policies to be identified with probability measures defined on a product
space with a fixed marginal at an input/state space (typically taken to be the Lebesgue measure in optimal determinis-
tic control) [227, 340], [74, Section 2.1], [322, p. 254], [223], [28, Theorem 2.2]. Thus, under the Young topology, one
associates with ΓR in (E.2) the probability measure induced on the product space X × U with a fixed marginal µ on X.
The generalization to stochastic control problems by considering more general input measures has been commonplace,
with applications also to partially observed stochastic control and decentralized stochastic control.
To appreciate the Young topology on control policies, we first present a relevant representation result (see Borkar [56]).
Let X, M be Borel spaces. Let P(X) denote the set of probability measures on X. Consider the set of probability measures

Θ := ζ ∈ P(X × M) :

f f
ζ(dx, dm) = P (dx) Q (dm|x), Q (·|x) = 1{f (x)∈·} , f : X → M (E.3)

on X×M with fixed input marginal P on X and with the stochastic kernel from X to M realized by any measurable function
f : X → M. We equip this set with the weak convergence topology. This set is the (Borel measurable) set of the extreme
336 E Relaxed Control Topologies

points of the set of probability measures on X × M with a fixed marginal P on X. For compact M, the Borel measurability
of Θ follows [251] since the set of probability measures on X × M with a fixed marginal P on X is a convex and compact
set in a complete separable metric space, and therefore, the set of its extreme points is Borel measurable; measurability for
the non-compact case follows from [56, Lemma 2.3]. Furthermore, given a fixed marginal P on X, any stochastic kernel Q
from X to M can almost surely be identified by a probability measure Ξ ∈ P(Θ) such that
Z
Q(·|x) = Ξ(dQf ) Qf (·|x). (E.4)
Θ

In particular, a randomized policy can thus be viewed as a mixture of deterministic policies.

Definition E.1.1 Convergence of Policies in ΓS under Young topology at reference (input) measure µ. Let µ be a σ-
finite measure. A sequence of stationary policies γn → γ ∈ ΓS at input µ if the joint measure (µγn ) → (µγ) weakly at
input P , i.e., for every continuous and bounded g : X × U → R with
Z
µ(dx) sup |g(x, u)| < ∞,
u∈U

Z   Z  
µ(dx) γn (du|x)g(x, u) → µ(dx) γ(du|x)g(x, u) (E.5)

With the above, we observe that the Young topology allows for a convex and compact formulation.
We also note that in the above, the reference measure does not need to be a probability measure (and we view weak
convergence to be one on signed measures that defines a locally convex space with (E.5) defining the semi-norms).
We finally note that since the marginal of the joint measure (µγn ) on X is fixed, the convergence in (E.5) is also in the
setwise-weak sense (with g(x, u) bounded but only continuous in u for every fixed x ∈ X, see Section D.4), following
Lemma D.4.1.

E.2 Borkar (Weak∗ ) Topology on Control Policies

In the stochastic setup, another topology is the one introduced by Borkar on relaxed controls [55] (see also [15, Section 2.4],
and [45] which [55] notes to be building on), formulated as a weak∗ topology on randomized policies viewed as maps from
states/measurements to the space of signed measures with bounded variation M(U) of which probability measures P(U)
is a subset. We also refer the reader to [104, 121] for further references on such a weak∗ formulation on relaxed controls,
in particular when instead of countably additive signed measures, finitely additive such measures are also considered.
Under the Borkar topology one studies ΓR in (E.2), with a weak∗ topology formulation, as a bounded subset of the set of
maps from X to the space of signed measures with finite variation viewed as the topological dual of continuous functions
vanishing at infinity, leading to a compact metric space by the Banach-Alaoglu theorem [133, Theorem 5.18] (and thus, as
the unit ball of L∞ (X, M(U)) = (L1 (X, C0 (U)))∗ is compact under the weak∗ topology, this leads to a compact metric
topology on relaxed control policies). We note that the presentations in [55, Section 3] and [15, Section 2.4] are slightly
different, though the induced topologies are identical. An equivalent representation of this topology is given in [15, Lemma
2.4.1] (see also [55, Lemma 3.1]).
See [12, 55, 256] for a detailed analysis on some implications in stochastic control theory in continuous time (such as
continuity of expected cost in control policies [55], approximation results [256] under various cost criteria, and continuity
of invariant measures of diffusions in control policies [12]).

Definition E.2.1 [55] [15, Lemma 2.4.1] Convergence of Policies in ΓS under Borkar topology. With X = Rd , a
sequence of stationary policies γn → γ ∈ ΓS in the Borkar topology if for every continuous and bounded g : X × U → R
and every f ∈ L1 (X) ∩ L2 (X)
E.3 Some Properties of Young and Borkar topologies 337
Z Z Z Z
f (x) γn (du|x)g(x, u)dx → f (x) γn (du|x)γ(du|x)g(x, u)dx (E.6)

Building on Lemma D.4.1, the functions g in Definition E.2.1 may be relaxed to be continuous only in u for every x ∈ X.
While X was taken to be Rn in [55], Saldi [268] generalized this to setups where X is a general standard Borel space with
a fixed input (probability) measure. The generalization by Saldi [268] is the following, where the input space X is arbitrary
standard Borel, though with a fixed input measure µ: Let C0 (U) be the Banach space of all continuous real functions on U
vanishing at infinity,
 endowed with the norm ∥g∥∞ = supu∈U |g(u)|. [268] formulated this topology via noting that with
L1 µ, C0 (U) denoting the set of all Bochner-integrable functions from X to C0 (U) endowed with the norm
Z
∥f ∥1 := ∥f (x)∥∞ µ(dx),
X


 
 C0 (U)
using the fact that  = M(U), and that the topological dual of L1 µ, C0 (U , ∥ · ∥1 can be identified with
L∞ µ, M(U) , ∥ · ∥∞ [76, Theorem 1.5.5, p. 27] (see also [104, 121] for further context on such duality results, in
particular when instead of countably additive signed measures, finitely additive such measures are considered); that is,
∗ 
L1 µ, C0 (U) = L∞ µ, M(U) .

E.3 Some Properties of Young and Borkar topologies

Lemma E.3.1 [347] Let η ≪ κ, where κ is a σ-finite and η is a finite measure. Then, γn → γ under Young topology at
input κ implies γn → γ (under Young topology) at input η.

Theorem E.3.1 [347] Let X = Rn and λ be the Lebesgue measure. Consider convergence in Young topology at some
input probability measure ψ.
(i) If ψ ≪ λ with h(x) = dψ
dλ (x) is positive everywhere, then convergence in Young topology at input measure ψ implies
convergence in Borkar topology.
(ii) If ψ ≪ λ, then convergence in Borkar topology implies convergence in Young topology at input ψ.

In [347], several results on the significance of these topologies on existence of optimal policies and approximations (on
near optimality of continuous policies or quantized policies in both measurement and action) have been presented.
References

1. J. Abounadi, D. Bertsekas, and V.S. Borkar. Learning algorithms for Markov decision processes with average cost. SIAM
Journal on Control and Optimization, 40(3):681–698, 2001.

2. M. Aicardi, F. Davoli, and R. Minciardi. Decentralized optimal control of Markov chains with a common past information set.
IEEE Transactions on Automatic Control, 32:1028–1031, November 1987.

3. C.D. Aliprantis and K.C. Border. Infinite Dimensional Analysis. Berlin, Springer, 3rd ed., 2006.

4. A. Almudevar. A stochastic contraction mapping theorem. Systems & Control Letters, 174:105482, 2023.

5. A. Almudevar. and E. F. Arruda. Optimal approximation schedules for a class of iterative algorithms, with an application to
multigrid value iteration. IEEE Transactions on Automatic Control, 57:3132–3146, 2012.

6. E. Altman. Constrained Markov Decision Processes. Chapman & Hall/CRC, Boca Raton, FL, 1999.

7. M. Andersland and D. Teneketzis. Information structures, causality, and non-sequential stochastic control, I: design-
independent properties. SIAM J. Control and Optimization, 30:1447 – 1475, 1992.

8. M. Andersland and D. Teneketzis. Information structures, causality, and non-sequential stochastic control, II: design-dependent
properties. SIAM J. Control and Optimization, 32:1726 – 1751, 1994.

9. B. D. O. Anderson and J. B. Moore. Optimal filtering. Englewood Cliffs, 21, 1979.

10. B. D. O. Anderson and J. B. Moore. Optimal control: Linear quadratic methods. Courier Corporation, 2007.

11. A. Arapostathis and V. S. Borkar. Uniform recurrence properties of controlled diffusions and applications to optimal control.
SIAM Journal on Control and Optimization, 48(7):4181–4223, 2010.

12. A. Arapostathis and V. S. Borkar. Average cost optimal control under weak ergodicity hypotheses: Relative value iterations.
Arxiv preprints, 1902.01048, 2019.

13. A. Arapostathis, V. S. Borkar, E. Fernandez-Gaucherand, M. K. Ghosh, and S. I. Marcus. Discrete-time controlled Markov
processes with average cost criterion: A survey. SIAM J. Control and Optimization, 31:282–344, 1993.

14. A. Arapostathis, V. S. Borkar, and M. K. Ghosh. Ergodic Control of Diffusion Processes, volume 143. Cambridge University
Press, 2012.

15. A. Arapostathis and S. Yüksel. Convex analytic method revisited: Further optimality results and performance of deterministic
policies in average cost stochastic control. Journal of Mathematical Analysis and Applications, 517(2):126567, 2023.

16. L. Arnold and W. Kliemann. On unique ergodicity for degenerate diffusions. Stochastics: an international journal of probability
and stochastic processes, 21(1):41–61, 1987.

17. E. F. Arruda, F. Ourique, J. Lacombe, and A. Almudevar. Accelerating the convergence of value iteration by using partial
transition functions. European Journal of Operational Research, 229:190–198, 2013.

18. M. Athans. Survey of decentralized control methods. Washington D.C, 1974. 3rd NBER/FRB Workshop on Stochastic Control.
340 References

19. K. B. Athreya and P. Ney. A new approach to the limit theory of recurrent Markov chains. Transactions of the American
Mathematical Society, 245:493–501, 1978.

20. R. J. Aumann. Mixed and behavior strategies in infinite extensive games. Technical report, Princeton University NJ, 1961.

21. R. J. Aumann. Agreeing to disagree. Annals of Statistics, 4:1236 – 1239, 1976.

22. R. J. Aumann. Correlated equilibrium as an expression of Bayesian rationality. Econometrica: Journal of the Econometric
Society, pages 1–18, 1987.

23. J. Backhoff-Veraguas, D. Bartl, M. Beiglböck, and M. Eder. Adapted Wasserstein distances and stability in mathematical
finance. Finance and Stochastics, 24(3):601–632, 2020.

24. A. Bain and D. Crisan. Fundamentals of stochastic filtering, volume 3. Springer, 2009.

25. W. L. Baker. Learning via Stochastic Approximation in Function Space. PhD Dissertation, Harvard University, Cambridge,
MA, 1997.

26. E. J. Balder. On ws-convergence of product measures. Mathematics of Operations Research, 26(3):494–518, 2001.

27. E.J. Balder. Generalized equilibrium results for games with incomplete information. Mathematics of Operations Research,
13(2):265–276, 1988.

28. Y. Bar-Shalom and E. Tse. Dual effect, certainty equivalence, and separation in stochastic control. IEEE Transactions on
Automatic Control, 19(5):494–500, October 1974.

29. Y. Bar-Shalom and E. Tse. Dual effect certainty equivalence and separation in stochastic control. IEEE Transactions on
Automatic Control, 19:494–500, October 1974.

30. E. Bayraktar, Yan Y. Dolinsky, and J. Guo. Continuity of utility maximization under weak convergence. Mathematics and
Financial Economics, pages 1–33, 2020.

31. M. Beiglböck and D. Lacker. Denseness of adapted processes among causal couplings. arXiv, pages arXiv–1805, 2018.

32. V. E. Beneš. Existence of optimal stochastic control laws. SIAM Journal on Control, 9(3):446–472, 1971.

33. A. Benveniste, M. Métivier, and P. Priouret. Adaptive algorithms and stochastic approximations, volume 22. Springer Science
& Business Media, 2012.

34. D. Bertsekas. Dynamic Programming and Optimal Control Vol. 1. Athena Scientific, 2000.

35. D. P. Bertsekas. Dynamic Programming and Stochastic Optimal Control. Academic Press, New York, New York, 1976.

36. D. P. Bertsekas and S. Shreve. Stochastic Optimal Control: The Discrete Time Case. Academic Press, New York, 1978.

37. D. P. Bertsekas and J. N. Tsitsiklis. Neuro-Dynamic Programming. Athena Scientific, Belmont, MA, 1996.

38. D.P. Bertsekas. Convergence of discretization procedures in dynamic programming. IEEE Trans. Autom. Control, 20(3):415–
419, Jun. 1975.

39. D.P Bertsekas. A new value iteration method for the average cost dynamic programming problem. SIAM journal on control
and optimization, 36(2):742–759, 1998.

40. D.P. Bertsekas and J.N. Tsitsiklis. Neuro-dynamic programming. Athena Scientific, 1996.

41. P. Billingsley. Convergence of Probability Measures. Wiley, New York, 1968.

42. P. Billingsley. Probability and Measure. Wiley, New York, 1995.

43. P. Billingsley. Probability and Measure. Wiley, 3rd edition, 1995.

44. J.-M. Bismut. Théorie Probabiliste du Contrôle des Diffusions, volume 181. 1973.

45. J.-M. Bismut. Partially observed diffusions and their control. SIAM Journal on Control and Optimization, 20(2):302–309,
1982.

46. D. Blackwell. Discrete dynamic programming. The Annals of Mathematical Statistics, pages 719–726, 1962.
References 341

47. D. Blackwell. Memoryless strategies in finite-stage dynamic programming. Annals of Mathematical Statistics, 35:863–865,
1964.

48. D. Blackwell. The stochastic processes of Borel gambling and dynamic programming. Annals of Statistics, pages 370–374,
1976.

49. D. Blackwell and L. Dubins. Merging of opinions with increasing information. Annals of Mathematical Statistics, 33:882–887,
1962.

50. D. Blackwell and C. Ryll-Nadrzewski. Non-existence of everywhere proper conditional distributions. Annals of Mathematical
Statistics, 34:223–225, 1963.

51. R. K. Boel, M. R. James, and I. R. Petersen. Robustness and risk-sensitive filtering. IEEE Transactions on Automatic Control,
47(3):451–461, 2002.

52. V. I. Bogachev. Measure Theory. Springer-Verlag, Berlin, 2007.

53. V. S. Borkar. The probabilistic structure of controlled diffusion processes. Acta Applicandae Mathematica, 11(1):19–48, 1988.

54. V. S. Borkar. A topology for Markov controls. Applied Mathematics and Optimization, 20(1):55–62, 1989.

55. V. S. Borkar. White-noise representations in stochastic realization theory. SIAM J. on Control and Optimization, 31:1093–1102,
1993.

56. V. S. Borkar. Probability Theory: An Advanced Course. Springer, New York, 1995.

57. V. S. Borkar. Average cost dynamic programming equations for controlled Markov chains with partial observations. SIAM J.
Control Optim., 39(3):673–681, 2000.

58. V. S. Borkar. Convex analytic methods in Markov decision processes. In Handbook of Markov Decision Processes, E. A.
Feinberg, A. Shwartz (Eds.), pages 347–375. Kluwer, Boston, MA, 2001.

59. V. S. Borkar. Dynamic programming for ergodic control with partial observations. Stochastic Processes and their Applications,
103:293–310, 2003.

60. V. S. Borkar. Ergodic control of diffusion processes. In Proceedings ICM, 2006.

61. V. S. Borkar. Dynamic programming for ergodic control of Markov chains under partial observations: A correction. SIAM J.
Control Optim., 45(6):2299–2304, 2007.

62. V. S. Borkar and A. Budhiraja. A further remark on dynamic programming for partially observed Markov processes. Stochastic
Processes and their Applications, 112:79–93, 2004.

63. V. S. Borkar and S. P. Meyn. The ODE method for convergence of stochastic approximation and reinforcement learning. SIAM
J. Control and Optimization, pages 447–469, December 2000.

64. V. S. Borkar and P. Varaiya. Asymptotic agreement in distributed estimation. IEEE Transactions Automatic Cotrol, 27:650–655,
June 1982.

65. Vivek S Borkar and Mrinal K Ghosh. Ergodic control of multidimensional diffusions i: The existence results. SIAM Journal
on Control and Optimization, 26(1):112–126, 1988.

66. V.S. Borkar. Learning algorithms for risk-sensitive control. In Proceedings of the 19th International Symposium on Mathemat-
ical Theory of Networks and Systems–MTNS, volume 5, 2010.

67. P. Bougerol. Kalman filtering with random coefficients and contractions. SIAM Journal on Control and Optimization,
31(4):942–959, 1993.

68. S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press, 2004.

69. A. Brandenburger and E. Dekel. Common knowledge with probability 1. J. Mathematical Economics, 16:237–245, 1987.

70. R. W. Brockett. Asymptotic stability and feedback stabilization. Differential Geometric Control Theory, 27(1):181–191, 1983.

71. A. Budhiraja. On invariant measures of discrete time filters in the correlated signal-noise case. The Annals of Applied Proba-
bility, 12(3):1096–1113, 2002.
342 References

72. P. E. Caines. Linear Stochastic Systems. John Wiley & Sons, New York, NY, 1988.

73. C. Castaing, P. R. De Fitte, and M. Valadier. Young measures on topological spaces: with applications in control theory and
probability theory, volume 571. Springer Science & Business Media, 2004.

74. [Link] and N.U. Ahmed. Dynamic team theory of stochastic differential decision systems with decentralized noisy
information structures via girsanov’s measure transformation. arXiv, abs/1309.1913, 2013.

75. P. Cembranos and J. Mendoza. Banach Spaces of Vector-Valued Functions. Springer-Verlag, 1997.

76. S. Chandak, V.S. Borkar, and P. Dodhia. Reinforcement learning in non-markovian environments. Systems & Control Letters,
185:105751, 2024.

77. H. S. Chang and S. I. Marcus. Approximate receding horizon approach for markov decision processes: average reward case.
Journal of Mathematical Analysis and Applications, 286(2):636–651, 2003.

78. C. D. Charalambous. Decentralized optimality conditions of stochastic differential decision problems via Girsanov’s measure
transformation. Mathematics of Control, Signals, and Systems, 28(3):1–55, 2016.

79. C. D. Charalambous and N. U. Ahmed. Equivalence of decentralized stochastic dynamic decision systems via Girsanov’s
measure transformation. In IEEE Conference on Decision and Control (CDC), pages 439–444. IEEE, 2014.

80. C. D. Charalambous and N. U. Ahmed. Maximum principle for decentralized stochastic differential decision systems. In IEEE
Conference on Decision and Control (CDC), pages 1846–1851. IEEE, 2014.

81. C. T. Chen. Linear Systems Theory and Design. Oxford University Press, Oxford, 1999.

82. T.S. Chiang, C.R. Hwang, and S.J. Sheu. Diffusion for global optimization in Rn . SIAM Journal on Control and Optimization,
25(3):737–753, 1987.

83. P. Chigansky and R. Liptser. On a role of predictor in the filtering stability. Electronic Communications in Probability,
11:129–140, 2006.

84. P. Chigansky, R. Liptser, and R. van Handel. Intrinsic methods in filter stability. Handbook of Nonlinear Filtering, 2009.

85. C. Y. Chong and M. Athans. On the periodic coordination of linear stochastic systems. Automatica, 12:321–335, 1976.

86. C.S. Chow and J. N. Tsitsiklis. An optimal one-way multigrid algorithm for discrete-time stochastic control. IEEE transactions
on automatic control, 36(8):898–914, 1991.

87. S. B. Connor and G. Fort. State-dependent Foster-Lyapunov criteria for subgeometric convergence of Markov chains. Stoch.
Process Appl., 119:176–4193, 2009.

88. O. Costa and F. Dufour. A sufficient condition for the existence of an invariant probability measure for markov processes.
Journal of Applied probability, 42(3):873–878, 2005.

89. O. Costa and F. Dufour. Average control of Markov decision processes with Feller transition probabilities and general action
spaces. Journal of Mathematical Analysis and Applications, 396(1):58–69, 2012.

90. L. Cregg, T. Linder, and S. Yüksel. Reinforcement learning for near-optimal design of zero-delay codes for markov sources.
IEEE Transactions on Information Theory, arXiv:2311.12609, 2024.

91. D. Crisan and A. Doucet. A survey of convergence results on particle filtering methods for practitioners. IEEE Transactions
on Signal Processing, 50(3):736–746, 2002.

92. I. Csiszár. Information-type measures of difference of probability distributions and indirect observation. studia scientiarum
Mathematicarum Hungarica, 2:229–318, 1967.

93. M. H. A Davis and P. Varaiya. Information states for linear stochastic systems. Journal of Mathematical Analysis and Appli-
cations, 37(2):384–402, 1972.

94. M. H. A Davis and P. Varaiya. Dynamic programming conditions for partially observable stochastic systems. SIAM Journal
on Control, 11(2):226–261, 1973.

95. Eugenio Della Vecchia, Silvia Di Marco, and Alain Jean-Marie. Illustrated review of convergence conditions of the value
iteration algorithm and the rolling horizon procedure for average-cost mdps. Annals of Operations Research, 199(1):193–214,
2012.
References 343

96. A. Dembo and O. Zeitouni. Large deviations techniques and applications, volume 38. Springer, 2010.

97. Y.E. Demirci, A. D. Kara, and S. Yüksel. Wasserstein regularity of nonlinear filters as belief-mdps, and implications on
ergodicity, optimality and learning for pomdps. In 2025 American Control Conference (ACC). IEEE, 2025.

98. Y.E. Demirci, A.D. Kara, and S. Yüksel. Average cost optimality of partially observed mdps: Contraction of non-linear filters
and existence of optimal solutions. SIAM Journal on Control and Optimization, 62:2859–2883, 2004.

99. Y.E. Demirci, A.D. Kara, and S. Yüksel. Refined bounds on near optimality finite window policies in pomdps and their
reinforcement learning. arXiv, 2024.

100. Y.E. Demirci and S. Yüksel. Unique ergodicity of non-linear filters via reachability and uniform weak continuity. Electronic
Communications in Probability, 30:1–15, 2025.

101. C. Derman. Finite state Markovian decision processes. Academic Press, Inc., 1970.

102. M.S. Derpich and S. Yüksel. Dual effect, certainty equivalence, and separation revisited: A counterexample and a relaxed
characterization for optimality. IEEE Transactions on Automatic Control, 68(2):1259–1266, 2023.

103. L. Devroye and L. Györfi. Non-parametric Density Estimation: The L1 View. John Wiley, New York, 1985.

104. J. Dieudonné. Sur le théorème de lebesgue-nikodym (v). Canadian Journal of Mathematics, 3:129–139, 1951.

105. R.L. Dobrushin. Central limit theorem for nonstationary Markov chains. i. Theory of Probability & Its Applications, 1(1):65–
80, 1956.

106. S. Dong, B. van Roy, and Z. Zhou. Simple agent, complex environment: Efficient reinforcement learning with agent states.
The Journal of Machine Learning Research, 23(1):11627–11680, 2022.

107. R. Douc, G. Fort, E. Moulines, and P. Soulier. Practical drift conditions for subgeometric rates of convergence. Ann. Appl.
Probab, 14:1353–1377, 2004.

108. R. Douc, E. Moulines, P. Priouret, and P. Soulier. Markov chains. Springer, 2018.

109. L. Dubins and D. Freedman. Measurable sets of measures. Pacific J. Math., 14:1211–1222, 1964.

110. R. M. Dudley. Real Analysis and Probability. Cambridge University Press, Cambridge, 2nd edition, 2002.

111. F. Dufour and T. Prieto-Rumeau. Approximation of Markov decision processes with general state space. J. Math. Anal. Appl.,
388:1254–1267, 2012.

112. F. Dufour and T. Prieto-Rumeau. Approximation of average cost Markov decision processes using empirical distributions and
concentration inequalities. Stochastics, pages 1–35, 2014.

113. J. Dugundji. An extension of Tietze’s theorem. Pacific Journal of Mathematics, 1(3):353–367, 1951.

114. P. G. Dupuis, M. R. James, and I. Petersen. Robust properties of risk-sensitive control. Mathematics of Control, Signals and
Systems, 13(4):318–332, 2000.

115. R. Durrett. Probability: Theory and Examples, volume 3. Cambridge university press, 2010.

116. E. B. Dynkin and A. A. Yushkevich. Controlled Markov Processes, volume 235. Springer, 1979.

117. E. Erdoğan and G. N. Iyengar. Ambiguous chance constrained problems and robust optimization. Mathematical Programming,
107(1-2):37–61, 2005.

118. P. M. Esfahani and D. Kuhn. Data-driven distributionally robust optimization using the Wasserstein metric: Performance
guarantees and tractable reformulations. Mathematical Programming, pages 1–52, 2017.

119. S. N. Ethier and T. G. Kurtz. Markov Processes: Characterization and Convergence, volume 282. John Wiley & Sons, 2009.

120. E. Even-Dar and Y. Mansour. Learning rates for q-learning. Journal of Machine Learning Research, 5:1–25, 2004.

121. H.O. Fattorini. Existence theory and the maximum principle for relaxed infinite-dimensional optimal control problems. SIAM
Journal on Control and Optimization, 32(2):311–331, 1994.

122. E. A. Feinberg. Non-randomized Markov and semi-Markov strategies in dynamic programming. Theory of Probability & Its
Applications, pages 116–126, 1982.
344 References

123. E. A. Feinberg. On measurability and representation of strategic measures in Markov decision processes. Institute of Mathe-
matical Statistics Lecture Notes. Eds. T. S. Ferguson, L. S. Shapley, J. B. MacQueen, pages 29–43, 1996.

124. E. A. Feinberg, P. O. Kasyanov, and N. V. Zadoianchuk. Berge’s theorem for noncompact image sets. J. Math. Anal. Appl.,
397(1):255–259, 2013.

125. E.A. Feinberg and P.O. Kasyanov. Mdps with setwise continuous transition probabilities. Operations Research Letters,
49(5):734–740, 2021.

126. E.A. Feinberg and P.O. Kasyanov. Equivalent conditions for weak continuity of nonlinear filters. Systems & Control Letters,
173:105458, 2023.

127. E.A. Feinberg, P.O. Kasyanov, and M.Z. Zgurovsky. Partially observable total-cost Markov decision process with weakly
continuous transition probabilities. Mathematics of Operations Research, 41(2):656–681, 2016.

128. E.A. Feinberg, P.O. Kasyanov, and M.Z. Zgurovsky. Markov decision processes with incomplete information and semiuniform
feller transition probabilities. SIAM Journal on Control and Optimization, 60(4):2488–2513, 2022.

129. W. Feller. An Introduction to Probability Theory and Its Applications. John Wiley and Sons, New York, 1971.

130. W. H. Fleming and H. M. Soner. Controlled Markov processes and viscosity solutions, volume 25. Springer Science & Business
Media, 2006.

131. W.H. Fleming and E. Pardoux. Optimal control for partially observed diffusions. SIAM J. Control Optim., 20(2):261–285,
1982.

132. S. R. Foguel. Positive operators on c(x). Proceedings of the American Mathematical Society, pages 295–297, 1969.

133. G. B. Folland. Real Analysis: Modern Techniques and Their Applications. John Wiley and Sons, 1999.

134. F. Forges. An approach to communication equilibria. Econometrica: Journal of the Econometric Society, pages 1375–1385,
1986.

135. P. K. Friz and M. Hairer. A course on rough paths. Springer, 2020.

136. Peter K. Friz and Martin Hairer. A course on rough paths. Universitext. Springer, Cham, [2020] ©2020. With an introduction
to regularity structures, Second edition of [ 3289027].

137. C. Gaskett and A. Zelinsky D. Wettergreen. Q-learning in continuous state and action spaces. In Australasian joint conference
on artificial intelligence, pages 417–428. Springer, 1999.

138. A. Gattami, B. M. Bernhardsson, and A. Rantzer. Robust team decision theory. IEEE Transactions on Automatic Control,
57:794–798, March 2012.

139. J. Geanakoplos and H. M. Polemarchakis. We can’t disagree forever. J. Economic Theory, pages 192–200, 1982.

140. T. T. Georgiou and A. Lindquist. The separation principle in stochastic control, redux. IEEE Transactions on Automatic
Control, 58(10):2481–2494, 2013.

141. A. Gersho. Stochastic stability of delta modulation. Bell Syst. Tech. J, 51(4):821–841, 1972.

142. J. K. Ghosh and R. V. Ramamoorthi. Bayesian Nonparametrics. Springer, New York, 2003.

143. I. I. Gihman and A. V. Skorohod. Controlled Stochastic Processes. Springer Science & Business Media, 2012.

144. I. V. Girsanov. On transforming a certain class of stochastic processes by absolutely continuous substitution of measures.
Theory of Probability & Its Applications, 5(3):285–301, 1960.

145. F. Le Gland and N. Oudjane. Stability and uniform approximation of nonlinear filters using the Hilbert metric and application
to particle filters. The Annals of Applied Probability, 14(1):144–187, 2004.

146. E. Gordienko and O. Hernández-Lerma. Average cost Markov control processes with weighted norms: Existence of canonical
policies. Appl. Math., 23(2):199–218, 1995.

147. E. Gordienko, E. Lemus-Rodrı́guez, and R. Montes de Oca. Discounted cost optimality problem: stability with respect to weak
metrics. Mathematical Methods of Operations Research, 68(1):77–96, 2008.
References 345

148. E. Gordienko, E. Lemus-Rodrı́guez, and R. Montes de Oca. Average cost markov control processes: stability with respect to
the kantorovich metric. Mathematical Methods of Operations Research, 70:13–33, 2009.

149. A. Gosavi. Reinforcement learning for long-run average cost. European journal of operational research, 155(3):654–674,
2004.

150. G. Grimmett and D. Stirzaker. Probability and random processes. Oxford university press, 2020.

151. A. Gupta, S. Yüksel, T. Başar, and C. Langbort. On the existence of optimal policies for a class of static and sequential dynamic
teams. SIAM Journal on Control and Optimization, 53:1681–1712, 2015.

152. M. Hairer. Ergodic properties of Markov processes. Lecture Notes, University of Warwick, 2006.

153. M. Hairer. Convergence of markov processes. Lecture Notes, 2010.

154. M. Hairer. On malliavins proof of hörmanders theorem. Bulletin des sciences mathematiques, 135(6-7):650–666, 2011.

155. B. Hajek. Optimal control of two interacting service stations. IEEE transactions on automatic control, 29(6):491–499, 1984.

156. B. Hajek. Lecture notes: Communication network analysis. University of Illinois at Urbana-Champaign, 2006.

157. B. Hajek. Random Processes for Engineers. Cambridge University Press, 2015.

158. R. Van Handel. Discrete time nonlinear filters with informative observations are stable. Electron. Commun. Probab, 13:562–
575, 2008.

159. O. Hernández-Lerma. Adaptive Markov Control Processes. Springer-Verlag, 1989.

160. O. Hernández-Lerma. Existence of average optimal policies in markov control processes with strictly unbounded costs. Ky-
bernetika, 29(1):1–17, 1993.

161. O. Hernández-Lerma. Adaptive Markov control processes, volume 79. Springer Science & Business Media, 2012.

162. O. Hernández-Lerma, R. Montes de Oca, and R. Cavazos-Cadena. Recurrence conditions for markov decision processes with
borel state space: a survey. Annals of Operations Research, 28(1):29–46, 1991.

163. O. Hernández-Lerma and J. B. Lasserre. Error bounds for rolling horizon policies in discrete-time markov control processes.
IEEE Transactions on Automatic Control, 35(10):1118–1124, 1990.

164. O. Hernández-Lerma and J. B. Lasserre. Discrete-Time Markov Control Processes: Basic Optimality Criteria. Springer, 1996.

165. O. Hernández-Lerma and J. B. Lasserre. Further topics on discrete-time Markov control processes. Springer, 1999.

166. O. Hernández-Lerma and J. B. Lasserre. Markov Chains and Invariant Probabilities. Birkhäuser, Basel, 2003.

167. O. Hernández-Lerma and J. B. Lasserre. Markov Chains and Invariant Probabilities. Birkhäuser, Basel, 2003.

168. C. J. Himmelberg, T. Parthasarathy, and F. S. Van Vleck. Optimal plans for dynamic programming problems. Mathematics of
Operations Research, 1(4):390–394, 1976.

169. K. Hinderer. Lipschitz continuity of value functions in markovian decision processes. Mathematical Methods of Operations
Research, 62:3–22, 2005.

170. Y. C. Ho and K. C. Chu. Team decision theory and information structures in optimal control problems - part I. IEEE Transac-
tions on Automatic Control, 17:15–22, February 1972.

171. Y. C. Ho and K. C. Chu. On the equivalence of information structures in static and dynamic teams. IEEE Transactions on
Automatic Control, 18(2):187–188, 1973.

172. I. Hogeboom-Burr and S. Yüksel. Sequential stochastic control (single or multi-agent) problems nearly admit change of
measures with independent measurements. Applied Mathematics and Optimization, 2023.

173. J. Hunter and B. Nachtergaele. Applied Analysis. World Scientific, Singapore, 2005.

174. N. Ikeda and S. Watanabe. Stochastic differential equations and diffusion processes. Elsevier, 2014.

175. O. C. Imer, S. Yüksel, and T. Başar. Optimal control of LTI systems over unreliable communication links. Automatica,
42(9):1429–1440, 2006.
346 References

176. G.N. Iyengar. Robust dynamic programming. Mathematics of Operations Research, 30(2):257–280, 2005.

177. T. Jaakkola, M.I. Jordan, and S.P. Singh. On the convergence of stochastic iterative dynamic programming algorithms. Neural
computation, 6(6):1185–1201, 1994.

178. D. Jacobson. Optimal stochastic linear systems with exponential performance criteria and their relation to deterministic differ-
ential games. IEEE Transactions on Automatic control, 18(2):124–131, 1973.

179. S. Jafarpour and A.D Lewis. Locally convex topologies and control theory. Mathematics of Control, Signals, and Systems,
28(4):29, 2016.

180. O. Kallenberg. Foundations of Modern Probability. Springer-Verlag, New York, 1997.

181. A. D. Kara, M. Raginsky, and S. Yüksel. Robustness to incorrect models and data-driven learning in average-cost optimal
stochastic control. Automatica, 139:110179, 2022.

182. A. D. Kara and S. Yüksel. Robustness to approximations and model learning in MDPs and POMDPs. In A. B. Piunovskiy
and Y. Zhang, editors, Modern Trends in Controlled Stochastic Processes: Theory and Applications, Volume III. Luniver Press,
2021.

183. A.D Kara, N. Saldi, and S. Yüksel. Weak Feller property of non-linear filters. Systems & Control Letters, 134:104–512, 2019.

184. A.D Kara, N. Saldi, and S. Yüksel. Q-learning for MDPs with general spaces: Convergence and near optimality via quantization
under weak continuity. Journal of Machine Learning Research, pages 1–34, 2023.

185. A.D Kara and S. Yüksel. Robustness to incorrect priors in partially observed stochastic control. SIAM Journal on Control and
Optimization, 57(3):1929–1964, 2019.

186. A.D Kara and S. Yüksel. Robustness to incorrect system models in stochastic control. SIAM Journal on Control and Optimiza-
tion, 58(2):1144–1182, 2020.

187. A.D Kara and S. Yüksel. Near optimality of finite memory feedback policies in partially observed markov decision processes.
Journal of Machine Learning Research, 23(11):1–46, 2022.

188. A.D Kara and S. Yüksel. Convergence of finite memory Q-learning for POMDPs and near optimality of learned policies under
filter stability. Mathematics of Operations Research, 48(4):2066–2093, 2023.

189. A.D. Kara and S. Yüksel. Q-learning for continuous state and action mdps under average cost criteria. arXiv preprint
arXiv:2308.07591, 2023.

190. A.D. Kara and S. Yüksel. Q-learning for stochastic control under general information structures and non-markovian environ-
ments. Transactions on Machine Learning Research, 2024. Featured Certification.

191. A.D Kara and S. Yüksel. Robustness to incorrect system models in stochastic control and application to data-driven learning.
In 2018 IEEE Conference on Decision and Control (CDC), pages 2753–2758, Dec 2018.

192. R. Khasminskii. Stochastic stability of differential equations. Springer, 2011.

193. W. Kliemann. Recurrence and invariant measures for degenerate diffusions. The Annals of Probability, 15(2):690–707, 1987.

194. T. Komorowski, S. Peszat, and T. Szarek. On ergodicity of some markov processes. 2010.

195. J.C. Krainak, J.L. Speyer, and S.I. Marcus. Static team problems – part I: Sufficient conditions and the exponential cost
criterion. IEEE Transactions on Automatic Control, 27:839–848, April 1982.

196. V. Krishnamurthy. Partially Observed Markov Decision Processes: Filtering, Learning and Controlled Sensing. Cambridge
University Press, 2 edition, 2025.

197. N.V. Krylov. Controlled Diffusion Processes, volume 14. Springer, 2008.

198. H. Kuhn. Extensive games and the problem of information. In Contributions to the Theory of Games, (H. Kuhn and A. Tucker,
editors), pages 193–216, 1953.

199. P. R. Kumar and P. Varaiya. Stochastic systems: Estimation, identification, and adaptive control. SIAM, 2015.

200. M. Kurano. The existence of a minimum pair of state and policy for markov decision processes under the hypothesis of doeblin.
SIAM journal on control and optimization, 27(2):296–307, 1989.
References 347

201. K. Kuratowski and C. Ryll-Nardzewski. A general theorem on selectors. Bull. Acad. Polon. Sci. Ser. Sci. Math. Astronom.
Phys, 13(1):397–403, 1965.

202. H. J. Kushner. Numerical methods for stochastic control problems in continuous time. SIAM Journal on Control and Opti-
mization, 28(5):999–1048, 1990.

203. H. J. Kushner. A partial history of the early development of continuous-time nonlinear stochastic systems theory. Automatica,
50(2):303–334, 2014.

204. H. J. Kushner and P. G. Dupuis. Numerical Methods for Stochastic Control Problems in Continuous Time, volume 24. Springer
Science & Business Media, 2001.

205. H.J. Kushner. Stochastic stability and control. Academic Press, New York, 1967.

206. H.J. Kushner. Introduction to Stochastic Control Theory. Holt, Rinehart and Winston, New York, 1972.

207. H.J. Kushner. Weak convergence methods and singularly perturbed stochastic control and filtering problems. Springer Science
& Business Media, 2012.

208. H.J. Kushner and G. Yin. Stochastic approximation and recursive algorithms and applications, 2003.

209. D. Lacker. Probabilistic compactification methods for stochastic optimal control and mean field games. 2018.

210. H. Lam. Robust sensitivity analysis for stochastic systems. Mathematics of Operations Research, 41(4):1248–1275, 2016.

211. H.J. Langen. Convergence of dynamic programming models. Mathematics of Operations Research, 6(4):493–512, Nov. 1981.

212. J. B. Lasserre. Invariant probabilities for Markov chains on a metric space. Statistics and Probability Letters, 34:259–265,
1997.

213. J. B. Lasserre. Sample-path average optimality for markov control processes. IEEE Transactions on Automatic Control,
44(10):1966–1971, 1999.

214. H. Lee and Y. Lim. Invariant metrics, contractions and nonlinear matrix equations. Nonlinearity, 21(4):857, 2008.

215. E. Lehrer, D. Rosenberg, and E. Shmaya. Signaling and mediation in games with common interests. Games and Economic
Behavior, 68:670–682, 2010.

216. B. C. Levy and M. Zorzi. A contraction analysis of the convergence of risk-sensitive filters. SIAM Journal on Control and
Optimization, 54(4):2154–2173, 2016.

217. D. Liberzon. Calculus of variations and optimal control theory: a concise introduction. Princeton university press, 2011.

218. A. Lindquist. On feedback control of linear stochastic systems. SIAM Journal on Control, 11(2):323–343, 1973.

219. D.G. Luenberger. Optimization by Vector Space Methods. John Wiley & Sons, New York, NY, 1969.

220. T. J. Lyons. Differential equations driven by rough signals. Revista Matemática Iberoamericana, 14(2):215–310, 1998.

221. A. S. Manne. Linear programming and sequential decision. Management Science, 6:259–267, April 1960.

222. X. Mao. Stochastic differential equations and applications. Elsevier, 2007.

223. E. Mascolo and L. Migliaccio. Relaxation methods in control theory. Applied Mathematics and Optimization, 20(1):97–103,
1989.

224. C. McDonald and S. Yüksel. Exponential filter stability via Dobrushin’s coefficient. Electronic Communications in Probability,
25, 2020.

225. C. McDonald and S. Yüksel. Robustness to incorrect priors and controlled filter stability in partially observed stochastic
control. SIAM Journal on Control and Optimization, 60(2):842–870, 2022.

226. C. McDonald and S. Yüksel. Stochastic observability and filter stability under several criteria. IEEE Transactions on Automatic
Control, 69(5):2931–2946, 2024.

227. E. J. McShane. Relaxed controls and variational problems. SIAM Journal on Control, 5(3):438–485, 1967.
348 References

228. F. C. Melo, S. P. Meyn, and I. M. Ribeiro. An analysis of reinforcement learning with function approximation. In Proceedings
of the 25th international conference on Machine learning, pages 664–671, 2008.

229. J.-F. Mertens, S. Sorin, and S. Zamir. Repeated games, volume 55. Cambridge University Press, 2015.

230. S. Meyn. Control systems and reinforcement learning. Cambridge University Press, 2022.

231. S. P. Meyn. Control Techniques for Complex Networks. Cambridge University Press, 2007.

232. S. P. Meyn and R. Tweedie. Markov Chains and Stochastic Stability. Springer-Verlag, London, 1993.

233. S. P. Meyn and R. Tweedie. State-dependent criteria for convergence of Markov chains. Ann. Appl. Prob, 4:149–168, 1994.

234. S. P. Meyn and R. L. Tweedie. Stability of Markovian processes ii: Continuous-time processes and sampled chains. Advances
in Applied Probability, 25(3):487–517, 1993.

235. S. P. Meyn and R. L. Tweedie. Stability of Markovian processes iii: Foster-lyapunov criteria for continuous-time processes.
Advances in Applied Probability, pages 518–548, 1993.

236. S.P. Meyn and R.L. Tweedie. State-dependent criteria for convergence of markov chains. Annals Appl. Prob., 4(1):149–168,
1994.

237. P. R. Milgrom and R. J. Weber. Distributional strategies for games with incomplete information. Mathematics of operations
research, 10(4):619–632, 1985.

238. S. K. Mitter. Nonlinear filtering of diffusion processes a guided tour. In Advances in Filtering and Optimal Stochastic Control,
pages 256–266. Springer, 1982.

239. O. Mrani-Zentar and S. Yüksel. Centralized reduction of decentralized stochastic control models and their weak-feller regular-
ity. arXiv:2408.13828, 2024.

240. A. Müller. How does the value function of a markov decision process depend on the transition probabilities? Mathematics of
Operations Research, 22(4):872–885, 1997.

241. A. Nayyar, A. Mahajan, and D. Teneketzis. Optimal control strategies in delayed sharing information structures. IEEE Trans-
actions on Automatic Control, 56:1606–1620, 2011.

242. A. Nayyar, A. Mahajan, and D. Teneketzis. The common-information approach to decentralized stochastic control. In Infor-
mation and Control in Networks, Editors: G. Como, B. Bernhardsson, A. Rantzer. Springer, 2013.

243. J. Neveu. Discrete-parameter martingales. revised edition, 1975.

244. L. Nielsen. Common knowledge, communication and convergence of beliefs. Mathematical Social Sciences, 8:1–14, 1984.

245. A. Nilim and L. El Ghaoui. Robust control of markov decision processes with uncertain transition matrices. Operations
Research, 53(5):780–798, 2005.

246. E. Nummelin. A splitting technique for harris recurrent markov chains. Z. Wahrscheinlichkeitstheoric verw. Gebiete, 43:309–
318, 1978.

247. E. Nummelin. A splitting technique for harris recurrent markov chains. Z. Wahrscheinlichkeitstheoric verw. Gebiete, 43:309–
318, 1978.

248. B. Øksendal. Stochastic Differential Equations. Springer, Berlin, 2003.

249. K.R. Parthasarathy. Probability Measures on Metric Spaces. AMS Bookstore, 1967.

250. I. Petersen, M. R. James, and P. Dupuis. Minimax optimal control of stochastic uncertain systems with relative entropy
constraints. IEEE Transactions on Automatic Control, 45(3):398–412, 2000.

251. R.P. Phelps. Lectures on Choquet’s theorem. Van Nostrand, New York:, 1966.

252. J. Pitman and M. Yor. A guide to brownian motion and related stochastic processes. arXiv preprint arXiv:1802.09679, 2018.

253. A. B. Piunovskiy. Controlled random sequences: methods of convex analysis and problems with functional constraints. Russian
Mathematical Surveys, 53(6):1233–1293, 1998.
References 349

254. P. Dai Pra, L. Meneghini, and W. J. Runggaldier. Connections between stochastic control and dynamic games. Mathematics of
Control, Signals and Systems, 9(4):303–326, 1996.

255. S. Pradhan, Z. Selk, and S. Yüksel. Robustness of optimal controlled diffusions with near-brownian noise via rough paths
theory. arXiv:2310.09967, 2023.

256. S. Pradhan and S. Yüksel. Continuity of cost in Borkar control topology and implications on discrete space and time approxi-
mations for controlled diffusions under several criteria. Electronic Journal of Probability, 29:1–32, 2024.

257. S. Pradhan and S. Yüksel. Near optimality of discrete-time approximations for controlled mckean-vlasov diffusions and
interacting particle systems. arXiv preprint arXiv:2510.21208, 2025.

258. Somnath Pradhan and Serdar Yüksel. Controlled diffusions under full, partial, and decentralized information: Existence of
optimal policies and discrete-time approximations. SIAM Journal on Control and Optimization, 63(5):3674–3702, 2025.

259. G. Da Prato and J. Zabczyk. Ergodicity for infinite dimensional systems, volume 229. Cambridge University Press, 1996.

260. R. Radner. Team decision problems. Ann. Math. Statist., 33:857–881, 1962.

261. R. Radner. Team decision problems. Annals of Mathematical Statistics, 33:857–881, 1962.

262. D. Rhenius. Incomplete information in Markovian decision models. Ann. Statist., 2:1327–1334, 1974.

263. H. Robbins and D. Siegmund. A convergence theorem for non negative almost supermartingales and some applications. In
Optimizing methods in statistics, pages 233–257. Elsevier, 1971.

264. G.O. Roberts and J.S. Rosenthal. General state space markov chains and mcmc algorithms. Probability Survery, 1:20–71,
2004.

265. K. W. Ross. Randomized and past-dependent policies for Markov decision processes with multiple constraints. Operations
Research, 37:474–477, May 1989.

266. S. M. Ross. On the nonexistence of ϵ-optimal randomized stationary policies in average cost markov decision models. The
Annals of Mathematical Statistics, 42(5):1767–1768, 1971.

267. E. P. Ryan. On brockett’s condition for smooth stabilizability and its necessity in a context of nonsmooth feedback. SIAM
Journal on Control and Optimization, 32(6):1597–1604, 1994.

268. N. Saldi. A topology for team policies and existence of optimal team policies in stochastic team theory. IEEE Transactions on
Automatic Control, 65(1):310–317, 2020.

269. N. Saldi, T. Linder, and S. Yüksel. Asymptotic optimality and rates of convergence of quantized stationary policies in stochastic
control. IEEE Trans. Automatic Control, 60:553 –558, 2015.

270. N. Saldi, T. Linder, and S. Yüksel. Finite Approximations in Discrete-Time Stochastic Control: Quantized Models and Asymp-
totic Optimality. Springer, Cham, 2018.

271. N. Saldi and S. Yüksel. Geometry of information structures, strategic measures and associated control topologies. Probability
Surveys, 19:450–532, 2022.

272. N. Saldi, S. Yüksel, and T. Linder. Finite-state approximation of Markov decision processes with unbounded costs and Borel
spaces. In IEEE Conf. Decision Control, Osaka. Japan, December 2015.

273. N. Saldi, S. Yüksel, and T. Linder. Near optimality of quantized policies in stochastic control under weak continuity conditions.
Journal of Mathematical Analysis and Applications, 435(1):321–337, 2016.

274. N. Saldi, S. Yüksel, and T. Linder. Finite model approximations and asymptotic optimality of quantized policies in decentral-
ized stochastic control. IEEE Transactions on Automatic Control, 62:2360 – 2373, 2017.

275. N. Saldi, S. Yüksel, and T. Linder. On the asymptotic optimality of finite approximations to Markov decision processes with
Borel spaces. Mathematics of Operations Research, 42(4):945–978, 2017.

276. N. Saldi, S. Yüksel, and T. Linder. Finite model approximations for partially observed Markov decision processes with dis-
counted cost. IEEE Transactions on Automatic Control, 65, 2020.

277. S. Sanjari, T. Başar, and S. Yüksel. Isomorphism properties of optimality and equilibrium solutions under equivalent informa-
tion structure transformations: Stochastic dynamic games and teams. SIAM Journal on Control and Optimization, 61(5):3102–
350 References

3130, 2023.

278. S. Sanjari, N. Saldi, and S. Yüksel. Optimality of independently randomized symmetric policies for exchangeable stochastic
teams with infinitely many decision makers. Mathematics of Operations Research, 48(3):1254–1285, 2023.

279. S. Sanjari and S. Yüksel. Optimal solutions to infinite-player stochastic teams and mean-field teams. IEEE Transactions on
Automatic Control, 66(3):1071–1086, 2020.

280. S. Sanjari and S. Yüksel. Optimal policies for convex symmetric stochastic dynamic teams and their mean-field limit. SIAM
Journal on Control and Optimization, 59(2):777–804, 2021.

281. A. V. Savkin and I. R. Petersen. Robust control of uncertain systems with structured uncertainty. Journal of Mathematical
Systems, Estimation, and Control, 6(3):1–14, 1996.

282. M. Schäl. A selection theorem for optimization problems. Archiv der Mathematik, 25(1):219–224, 1974.

283. M. Schäl. Conditions for optimality in dynamic programming and for the limit of n-stage optimal policies to be optimal. Z.
Wahrscheinlichkeitsth, 32:179–296, 1975.

284. M. Schäl. On dynamic programming: compactness of the space of policies. Stochastic Processes and their Applications,
3(4):345–364, 1975.

285. M. Schäl. Average optimality in dynamic programming with general state space. Mathematics of operations Research,
18(1):163–172, 1993.

286. L. Schenato, B. Sinopoli, M. Franceschetti, K. Poolla, and S. S. Sastry. Foundations of control and estimation over lossy
networks. Proceedings of the IEEE, 95(1):163–187, 2007.

287. R. Serfozo. Convergence of Lebesgue integrals with varying measures. Sankhyā: The Indian Journal of Statistics, Series A,
pages 380–402, 1982.

288. S. Shreve and D. P. Bertsekas. Universally measurable policies in dynamic programming. Mathematics of Operations Research,
4(1):15–30, 1979.

289. S. P. Singh, T. Jaakkola, and M. I. Jordan. Reinforcement learning with soft state aggregation. Advances in neural information
processing systems, pages 361–368, 1995.

290. S.P. Singh, T. Jaakkola, and M.I. Jordan. Learning without state-estimation in partially observable markovian decision pro-
cesses. In Machine Learning Proceedings 1994, pages 284–292. Elsevier, 1994.

291. E. D. Sontag. Mathematical Control Theory: Deterministic Finite Dimensional Systems, volume 6. Springer Science &
Business Media, 2013.

292. R.B Sowers and A.M. Makowski. Discrete-time filtering for linear systems with non-Gaussian initial conditions: asymptotic
behavior of the difference between the MMSE and LMSE estimates. IEEE Transactions on Automatic Control, 37(1):114–120,
1992.

293. S. M. Srivastava. A course on Borel sets, volume 180. Springer Science & Business Media, 2008.

294. L. Stettner. On the existence and uniqueness of invariant measure for continuous time markov processes. Technical report,
BROWN UNIV PROVIDENCE RI LEFSCHETZ CENTER FOR DYNAMICAL SYSTEMS, 1986.

295. Richard H Stockbridge. Time-average control of martingale problems: Existence of a stationary solution. Annals of Probability,
pages 190–205, 1990.

296. C. Striebel. Sufficient statistics in the optimum control of stochastic systems. J. Math. Anal. Appl, 12:576–592, 1965.

297. D. Stroock and S. R. S. Varadhan. On the support of diffusion processes with applications to the strong maximum principle.
In Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability (Univ. California, Berkeley, Calif.,
1970/1971), volume 3, pages 333–359, 1972.

298. H. Sun and H. Xu. Convergence analysis for distributionally robust optimization and equilibrium problems. Mathematics of
Operations Research, 41:377–401, 07 2015.

299. C. Szepesvári. The asymptotic convergence-rate of q-learning. Advances in neural information processing systems, 10, 1997.

300. C. Szepesvári. Algorithms for Reinforcement Learning. Morgan and Claypool, 2010.
References 351

301. C. Szepesvári. Algorithms for reinforcement learning. volume 4, pages 1–103, 2010.

302. C. Szepesvàri and W.D. Smart. Interpolation-based q-learning. 2004.

303. C. Szepesvári and M.L. Littman. A unified analysis of value-function-based reinforcement-learning algorithms. Neural com-
putation, 11(8):2017–2060, 1999.

304. D. Teneketzis. On information structures and nonsequential stochastic control. CWI Quarterly, 9:241–260, 1996.

305. D. Teneketzis. On the structure of optimal real-time encoders and decoders in noisy communication. IEEE Transactions on
Information Theory, 52:4017–4035, September 2006.

306. J. N. Tsitsiklis. Asynchronous stochastic approximation and Q-learning. Machine Learning, 16:185–202, 1994.

307. J. N. Tsitsiklis and M. Athans. Convergence and asymptotic agreement in distributed decision problems. IEEE Transactions
on Automatic Control, pages 42–50, January 1984.

308. P. Tuominen and R.L. Tweedie. Subgeometric rates of convergence of f-ergodic markov chains. Adv. Annals Appl. Prob.,
26(3):775–798, September 1994.

309. R. L. Tweedie. Topological conditions enabling use of Harris methods in discrete and continuous time. Acta Appl. Math.,
34(1-2):175–188, 1994.

310. R. L. Tweedie. Drift conditions and invariant measures for Markov chains. Stochastic Processes and Their Applications,
92:345–354, 2001.

311. V. A. Ugrinovskii. Robust H-infinity control in the presence of stochastic uncertainty. International Journal of Control,
71(2):219–237, 1998.

312. R. van Handel. Stochastic calculus, filtering, and stochastic control. Course notes., URL [Link] princeton. edu/˜
rvan/acm217/ACM217. pdf, 2007.

313. R. van Handel. Observability and nonlinear filtering. Probability theory and related fields, 145(1-2):35–74, 2009.

314. R. van Handel. The stability of conditional Markov processes and Markov chains in random environments. Ann. Probab,
37:1876–1925, 2009.

315. R. van Handel. Nonlinear filtering and systems theory. In Proceedings of the 19th International Symposium on Mathematical
Theory of Networks and Systems (MTNS semi-plenary paper), 2010.

316. S.R.S. Varadhan. Probability theory, volume 7 of Courant Lecture Notes in Mathematics, volume 1. New York University
Courant Institute of Mathematical Sciences, New York.

317. O. Vega-Amaya. The average cost optimality equation: a fixed point approach. Bol. Soc. Mat. Mexicana, 9(3):185–195, 2003.

318. Oscar Vega-Amaya. The average cost optimality equation: a fixed point approach. Bol. Soc. Mat. Mexicana, 9(1):185–195,
2003.

319. M. Vidyasagar. Convergence of stochastic approximation via martingale and converse lyapunov methods. Mathematics of
Control, Signals, and Systems, pages 1–24, 2023.

320. C. Villani. Optimal Transport: Old and New. Springer, 2008.

321. J. C. Walrand and P. Varaiya. Optimal causal coding-decoding problems. IEEE Transactions on Information Theory, 19:814–
820, November 1983.

322. J. Warga. Optimal Control of Differential and Functional Equations. Academic press, 2014.

323. C. J. C. H. Watkins and P. Dayan. Q-learning. Machine Learning, 8:279–292, 1992.

324. G. L. Wise. A note on a common misconception in estimation. Systems and Cont. Letters, 5:355–356, April 1985.

325. H. S. Witsenhausen. A counterexample in stochastic optimal control. SIAM J. Contr., 6:131–147, 1968.

326. H. S. Witsenhausen. On information structures, feedback and causality. SIAM J. Control, 9:149–160, May 1971.

327. H. S. Witsenhausen. A standard form for sequential stochastic control. Mathematical Systems Theory, 7:5–11, 1973.
352 References

328. H. S. Witsenhausen. The intrinsic model for discrete stochastic control: Some open problems. Lecture Notes in Econ. and
Math. Syst., Springer-Verlag, 107:322–335, 1975.

329. H. S. Witsenhausen. The intrinsic model for discrete stochastic control: Some open problems. In Control Theory, Numerical
Methods and Computer System Modelling, pages 322–335, A. Bensoussan and J. L. Lions Springer-Verlag 107, 1975. Lecture
Notes in Economics and Mathematical Systems.

330. H. S. Witsenhausen. On the structure of real-time source coders. Bell Syst. Tech. J, 58:1437–1451, July/August 1979.

331. H. S. Witsenhausen. Equivalent stochastic control problems. Math. Control, Signals and Systems, 1:3–11, 1988.

332. E. Wong and B.E. Hajek. Stochastic Processes in Engineering Systems. Springer-Verlag, New York, 1985.

333. E. Wong and M. Zakai. On the convergence of ordinary integrals to stochastic integrals. The Annals of Mathematical Statistics,
36(5):1560–1564, 1965.

334. W. M. Wonham. On the separation theorem of stochastic control. SIAM Journal on Control, 6(2):312–326, 1968.

335. R.G. Wood, T. Linder, and S. Yüksel. Optimal zero delay coding of Markov sources: Stationary and finite memory codes.
IEEE Transactions on Information Theory, 63:5968–5980, 2017.

336. D. T. H. Worm and S. C. Hille. Ergodic decompositions associated with regular markov operators on polish spaces. Ergodic
Theory and Dynamical Systems, 31(2):571 – 597, 2010.

337. H. Xu and S. Mannor. Distributionally robust Markov decision processes. In J. D. Lafferty, C. K. I. Williams, J. Shawe-
Taylor, R. S. Zemel, and A. Culotta, editors, Advances in Neural Information Processing Systems 23, pages 2505–2513. Curran
Associates, Inc., 2010.

338. T. Yoshikawa. Dynamic programming approach to decentralized control problems. IEEE Transactions on Automatic Control,
20:796–797, 1975.

339. L. C. Young. An inequality of the hölder type, connected with stieltjes integration. Acta Mathematica, 67:251–282, 1936.

340. L.C. Young. Generalized curves and the existence of an attained absolute minimum in the calculus of variations. Comptes
Rendus de la Societe des Sci. et des Lettres de Varsovie, 30:212–234, 1937.

341. H. Yu. Some proof details for asynchronous stochastic approximation algorithms. 2012.

342. S. Yüksel. Stochastic nestedness and the belief sharing information pattern. IEEE Transactions on Automatic Control, 54:2773–
2786, December 2009.

343. S. Yüksel. On optimal causal coding of partially observed Markov sources in single and multi-terminal settings. IEEE Trans-
actions on Information Theory, 59:424–437, January 2013.

344. S. Yüksel. On stochastic stability of a class of non-Markovian processes and applications in quantization. SIAM J. on Control
and Optimization, 55:1241–1260, 2017.

345. S. Yüksel. A note on the separation of optimal quantization and control policies in networked control. SIAM Journal on Control
and Optimization, 57(1):773–782, 2019.

346. S. Yüksel. A universal dynamic program and refined existence results for decentralized stochastic control. SIAM Journal on
Control and Optimization, 58(5):2711–2739, 2020.

347. S. Yüksel. On Borkar and Young relaxed control topologies and continuous dependence of invariant measures on control
policy. SIAM Journal on Control and Optimization, 62(4):2367–2386, 2024.

348. S. Yüksel. Another look at partially observed optimal stochastic control: Existence, ergodicity, and approximations without
belief-reduction. Applied Mathematics & Optimization, 91(1):16, 2025.

349. S. Yüksel and T. Başar. Stochastic Networked Control Systems: Stabilization and Optimization under Information Constraints.
Springer, New York, 2013.

350. S. Yüksel and S. P. Meyn. Random-time, state-dependent stochastic drift for Markov chains and application to stochastic
stabilization over erasure channels. IEEE Transactions on Automatic Control, 58:47 – 59, January 2013.

351. S. Yüksel and N. Saldi. Convex analysis in decentralized stochastic control, strategic measures and optimal solutions. SIAM J.
on Control and Optimization, 55:1–28, 2017.
References 353

352. A.A. Yushkevich. Reduction of a controlled Markov model with incomplete data to a problem with complete information in
the case of Borel state and control spaces. Theory Prob. Appl., 21:153–158, 1976.

353. Y. Zhou, Y. Song, and S. Yüksel. Robustness to model approximation, learning, and sample complexity in wasserstein regular
mdps. arXiv preprint arXiv:2410.14116, 2024.

354. R. Zurkowski, S. Yüksel, and T. Linder. On rates of convergence for Markov chains under random time state-dependent
stochastic drift criteria. IEEE Transactions on Automatic Control, 61(1):145–155, 2015.

You might also like