Experts in a Markov Decision Process

Even-Dar, Eyal; Kakade, Sham; Mansour, Yishay

Experts in a Markov Decision Process

Files

2730_experts_in_a_markov_decision_process.pdf (73.63 KB)

Penn collection

Statistics Papers

Subject

Statistics and Probability

Permalink

https://repository.upenn.edu/handle/20.500.14332/47845

View all metadata

Author

Even-Dar, Eyal

Kakade, Sham

Mansour, Yishay

Abstract

We consider an MDP setting in which the reward function is allowed to change during each time step of play (possibly in an adversarial manner), yet the dynamics remain fixed. Similar to the experts setting, we address the question of how well can an agent do when compared to the reward achieved under the best stationary policy over time. We provide efficient algorithms, which have regret bounds with no dependence on the size of state space. Instead, these bounds depend only on a certain horizon time of the process and logarithmically on the number of actions. We also show that in the case that the dynamics change over time, the problem becomes computationally hard.

Date of presentation

2004-01-01

Conference name

Statistics Papers

Conference dates

2023-05-17T15:04:37.000

Collection

Presentations