Podcast Episodes

Back to Search
Q♯: Distributional RL for Optimal LLM Post-Training
Q♯: Distributional RL for Optimal LLM Post-Training

This podcast introduces Q♯, a novel reinforcement learning algorithm tailored for post-training large language models (LLMs) by utilizing distributio…

1 year, 5 months ago

Short Long
View Episode
Scaling Test-Time Compute Without Verification or RL is Suboptimal
Scaling Test-Time Compute Without Verification or RL is Suboptimal


The paper presents a theoretical analysis comparing verifier-based (VB) and verifier-free (VF) algorithms for training large language models (LLMs) u…

1 year, 5 months ago

Short Long
View Episode
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Longer version

1 year, 5 months ago

Short Long
View Episode
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

The paper optimizes test-time compute as a meta-reinforcement learning problem It emphasizes balancing exploration and exploitation to minimize cumul…

1 year, 5 months ago

Short Long
View Episode
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

The paper surveys limitations of reinforcement learning from human feedback (RLHF). It highlights challenges in training AI systems with RLHF. Propos…

1 year, 5 months ago

Short Long
View Episode
Revisiting Superficial Alignment Hypothesis
Revisiting Superficial Alignment Hypothesis

The paper revisits the Superficial Alignment Hypothesis. It studies post-training scaling behavior with finetuning examples. Performance scales as a …

1 year, 5 months ago

Short Long
View Episode
Diagnostic uncertainty: teaching language Models to describe open-ended uncertainty
Diagnostic uncertainty: teaching language Models to describe open-ended uncertainty

The paper introduces diagnostic uncertainty in language models.It enables models to describe their uncertainty openly.Improved accuracy and reduced e…

1 year, 5 months ago

Short Long
View Episode
Language Model Personalization via Reward Factorization
Language Model Personalization via Reward Factorization


The paper introduces a personalized framework for LLMs. It utilizes user-specific rewards from minimal feedback. The method achieves significant pers…

1 year, 5 months ago

Short Long
View Episode
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration


The paper explores efficient exploration techniques in language model alignment It introduces SpannerSampling for optimal data efficiency in reinforc…

1 year, 5 months ago

Short Long
View Episode
How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach


The paper studies reasoning length and model performance tradeoff. It explores compression strategies for large language models (LLMs). Token complex…

1 year, 5 months ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us