Podcast Episodes

Back to Search
Language Models Can Control Their Own Attention
Language Models Can Control Their Own Attention

Researchers have introduced Declarative Attention (DA), a protocol that enables large language models to autonomously manage their own focus during l…

15 hours ago

Short Long
View Episode
AI Finds A Way
AI Finds A Way

This paper introduces a comprehensive collection of anecdotes documenting instances where artificial intelligence systems developed innovative yet un…

1 day, 16 hours ago

Short Long
View Episode
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

This research paper investigates how feature entanglement in large language models prevents precise, localized interventions on specific concepts. Th…

2 days, 12 hours ago

Short Long
View Episode
TTPO: Test-Time Policy Optimization
TTPO: Test-Time Policy Optimization

This paper introduces Test-Time Policy Optimization (TTPO), a novel method for improving the mathematical reasoning of large language models without …

3 days, 6 hours ago

Short Long
View Episode
Demystifying Reinforcement Learning Post-Training of Language Models
Demystifying Reinforcement Learning Post-Training of Language Models

This paper deconstructs the mechanics of reinforcement learning (RL) post-training for large language models to determine how different factors influ…

4 days, 19 hours ago

Short Long
View Episode
Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses
Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses

Recuris is a recursive architectural framework designed to enhance the performance of large language model agents during complex, long-horizon tasks.…

5 days, 15 hours ago

Short Long
View Episode
TailSFT: Filtered Fine-Tuning Improves Post-Training Performance
TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

Researchers introduce TailSFT, a modified supervised fine-tuning algorithm designed to better prepare language models for subsequent reinforcement le…

6 days, 10 hours ago

Short Long
View Episode
SPADE: Self-Play in Adaptive Synthetic Executable Environments
SPADE: Self-Play in Adaptive Synthetic Executable Environments

This paper introduces SPADE, a reinforcement learning framework that enables a single large language model to achieve open-ended self-improvement by …

1 week ago

Short Long
View Episode
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

This paper introduces ACES (Agentic Continuous Evaluation of Skills), a comprehensive framework developed by NVIDIA to move beyond static document sc…

1 week, 3 days ago

Short Long
View Episode
Impression Share Prediction: An Offline Evaluation Task for Ranking Systems
Impression Share Prediction: An Offline Evaluation Task for Ranking Systems

Researchers from Meta Platforms propose a novel offline evaluation task called impression share prediction to better anticipate how new ranking model…

1 week, 4 days ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us