RIML Lab: post #255 — TG.ME

🤖 RL Journal Club

This Week's Presentation:

🔹 Title: Test Time Exploration to Achieve Generalization in Zero-Shot RL
🔸 Presenter: Alireza Farajtabrizi

🌀 Abstract:
This paper studies zero-shot generalization in reinforcement learning, where an agent is trained on a set of tasks but must perform well on unseen test environments. The authors argue that standard reward-maximizing RL agents can overfit to training tasks, especially in environments where simple invariance-based methods fail. Their key insight is that exploration behavior is harder to memorize than reward-seeking behavior and can therefore generalize better.

To build on this idea, the paper introduces Explore to Generalize (ExpGen), an algorithm that combines a maximum-entropy exploration policy with an ensemble of reward-seeking agents. At test time, when the ensemble agrees on an action, the agent exploits that decision; when the ensemble is uncertain, the agent switches to the exploration policy to reach new parts of the state space. Experiments on ProcGen show strong improvements on challenging tasks such as Maze and Heist, setting new state-of-the-art results in several zero-shot RL settings.

📄 Paper: Explore to Generalize in Zero-Shot RL (NeurIPS 2023)

Session Details:

* 📅 Date: Tuesday سه‌شنبه
* 🕒 Time: 15:30 - 16:30
* 🌐 Location: Online at vc.sharif.edu/ch/rohban (http://vc.sharif.edu/ch/rohban)

We look forward to your participation! ✌️
proceedings.neurips.cc/paper_files/paper/2023/file/c793577b644268259b1416464a6cdb8c-Paper-Conference.pdf
June 22, 2026 3K 26