Podcast Episodes
Back to Search
Reinforcement Learning in Non-Stationary Environments
This academic paper introduces Non-Stationary Natural Actor-Critic (NS-NAC), a novel model-free, policy-based reinforcement learning algorithm design…
1Â year, 1Â month ago
Personalized Policy Learning from Heterogeneous Data
This document introduces a novel framework for offline reinforcement learning (RL), focusing on optimizing individual policies when data comes from d…
1Â year, 1Â month ago
Boosting Reinforcement Learning with Human Feedback via SeRA
This article from Amazon Science, published in May 2025, focuses on machine learning and conversational AI, specifically addressing improvements in r…
1Â year, 1Â month ago
AXIOM: Active Inference Object-Centric World Models
This document introduces AXIOM, a novel artificial intelligence architecture designed to learn how to play games efficiently using object-centric mod…
1Â year, 1Â month ago
Entropy and Reinforcement Learning for LLMs
This academic paper explores a critical issue in reinforcement learning (RL) with large language models (LLMs): the rapid decline of policy entropy, …
1Â year, 1Â month ago
FLEX Robot-Agnostic Force-Based Manipulation Learning
AI and Robotics
1Â year, 1Â month ago
Agent RL Scaling for Mathematical Problem Solving
This academic paper explores ZeroTIR, a novel method for training Large Language Models (LLMs) to spontaneously use external tools, specifically Pyth…
1Â year, 1Â month ago
Beyond Reward: Limits of RL in LLM Reasoning
This academic paper critically re-evaluates the widespread belief that Reinforcement Learning with Verifiable Rewards (RLVR) enhances the fundamental…
1Â year, 1Â month ago
Reward Model Variance in RLHF
This document investigates how the quality of a reward model impacts the training efficiency of language models using Reinforcement Learning from Hum…
1Â year, 1Â month ago
Power Grid Topological Control with Graph Reinforcement Learning
This document presents research on improving power grid management through reinforcement learning. The authors introduce a model-free approach using …
1Â year, 1Â month ago