Podcast Episodes
Back to Search
Q♯: Distributional RL for Optimal LLM Post-Training
This podcast introduces Q♯, a novel reinforcement learning algorithm tailored for post-training large language models (LLMs) by utilizing distributio…
1 year, 5 months ago
Scaling Test-Time Compute Without Verification or RL is Suboptimal
1 year, 5 months ago
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Longer version
1 year, 5 months ago
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
The paper optimizes test-time compute as a meta-reinforcement learning problem It emphasizes balancing exploration and exploitation to minimize cumul…
1 year, 5 months ago
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
The paper surveys limitations of reinforcement learning from human feedback (RLHF). It highlights challenges in training AI systems with RLHF. Propos…
1 year, 5 months ago
Revisiting Superficial Alignment Hypothesis
The paper revisits the Superficial Alignment Hypothesis. It studies post-training scaling behavior with finetuning examples. Performance scales as a …
1 year, 5 months ago
Diagnostic uncertainty: teaching language Models to describe open-ended uncertainty
The paper introduces diagnostic uncertainty in language models.It enables models to describe their uncertainty openly.Improved accuracy and reduced e…
1 year, 5 months ago
Language Model Personalization via Reward Factorization
1 year, 5 months ago
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
1 year, 5 months ago
How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
1 year, 5 months ago