Skip to main content
The following code shows the estimation of the q value function for a policy, the optimal q_star and the optimal policy for the cleaning robot problem in the deterministic case.
Output from cell 2
Output from cell 2
Output from cell 2
Output from cell 2
Output from cell 2
Output from cell 2