Stochastic Control of Continuous-Time and Continuous-State Systems via Direct Comparison
Xi‐Ren Cao · Rare & Special e-Zone (The Hong Kong University of Science and Technology) · 2009
The standard approach to stochastic control is dy- namic programming. In this paper, we introduce an alternative approach based on direct comparison of the performance of any two policies, by modeling the state process as a continuous-time and continuous state Markov process. The approach provides a unified framework for stochastic control and other optimization theory and methodologies including Markov decision processes, perturbation analysis, and reinforcement learning. The new insights obtained may lead to new research topics. Control and performance optimization of stochastic sys- tems is a multi-disciplinary subject that has attracted wide attention from many research communities. Many different research areas have been established; these different areas take different perspectives and therefore have different mod- els of the systems and formulations of the problem. The traditional stochastic control theory deals with con- tinuous systems modeled by stochastic differential equations, and the standard approach is dynamic programming (1). The approach is particularly suitable for finite-horizon problems; it works backwards in time. The problem with infinite horizon can be treated as the limiting cases of the finite- horizon problem when time going to infinity, and the long- run average cost problem can be treated as the limiting case of the problems with discounted costs. In the approach, the Hamilton- Jacobi-Bellman (HJB) equation for the optimal policies is first established with the dynamic programming principle, a verification theorem is then proved which states that the solution to the HJB equation indeed provides the value function and from which an optimal control process can be constructed. The HJB equations are usually differ- ential equations, and the concept of viscosity solution is introduced when the value functions are not differentiable (8).