Conservative Offline Distributional Reinforcement Learning

Ma, Yecheng Jason; Jayaraman, Dinesh; Bastani, Osbert

Computer Science > Machine Learning

arXiv:2107.06106 (cs)

[Submitted on 12 Jul 2021 (v1), last revised 26 Oct 2021 (this version, v2)]

Title:Conservative Offline Distributional Reinforcement Learning

Authors:Yecheng Jason Ma, Dinesh Jayaraman, Osbert Bastani

View PDF

Abstract:Many reinforcement learning (RL) problems in practice are offline, learning purely from observational data. A key challenge is how to ensure the learned policy is safe, which requires quantifying the risk associated with different actions. In the online setting, distributional RL algorithms do so by learning the distribution over returns (i.e., cumulative rewards) instead of the expected return; beyond quantifying risk, they have also been shown to learn better representations for planning. We propose Conservative Offline Distributional Actor Critic (CODAC), an offline RL algorithm suitable for both risk-neutral and risk-averse domains. CODAC adapts distributional RL to the offline setting by penalizing the predicted quantiles of the return for out-of-distribution actions. We prove that CODAC learns a conservative return distribution -- in particular, for finite MDPs, CODAC converges to an uniform lower bound on the quantiles of the return distribution; our proof relies on a novel analysis of the distributional Bellman operator. In our experiments, on two challenging robot navigation tasks, CODAC successfully learns risk-averse policies using offline data collected purely from risk-neutral agents. Furthermore, CODAC is state-of-the-art on the D4RL MuJoCo benchmark in terms of both expected and risk-sensitive performance.

Comments:	NeurIPS 2021
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2107.06106 [cs.LG]
	(or arXiv:2107.06106v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2107.06106

Submission history

From: Yecheng Jason Ma [view email]
[v1] Mon, 12 Jul 2021 15:38:06 UTC (1,438 KB)
[v2] Tue, 26 Oct 2021 18:16:17 UTC (1,556 KB)

Computer Science > Machine Learning

Title:Conservative Offline Distributional Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Conservative Offline Distributional Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators