Podcast Episodes

Back to Search
Harness RL is Meta-Learning: Training to Self-Improve at Test Time
Harness RL is Meta-Learning: Training to Self-Improve at Test Time

This paper introduces harness RL, a novel meta-learning framework designed to enable large language models to self-improve during test-time adaptatio…

4 weeks, 2 days ago

Short Long
View Episode
Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models
Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models

This paper investigates a critical strategic mismatch between Large Language Models (LLMs) and human decision-makers in competitive environments. Thr…

1 month ago

Short Long
View Episode
When Does LeJEPA Learn a World Model?
When Does LeJEPA Learn a World Model?

This research paper introduces a mathematical framework to prove that LeJEPA (a specific self-supervised learning architecture) can accurately recove…

1 month ago

Short Long
View Episode
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems

This research paper investigates Role Drift, a failure mode in compound AI systems where individual modules abandon their specific instructions to fi…

1 month ago

Short Long
View Episode
Do you really need to pretrain Q-functions for online RL fine-tuning?
Do you really need to pretrain Q-functions for online RL fine-tuning?

Research from Stanford University challenges the conventional assumption that pre-training a Q-function on offline data improves reinforcement learni…

1 month, 1 week ago

Short Long
View Episode
The Evolution of Digital Search: From Blue Links to Delegated Decision-Making
The Evolution of Digital Search: From Blue Links to Delegated Decision-Making

Digital search is transitioning from a human-centered discovery process based on links and keywords to an agent-mediated system of delegated decision…

1 month, 1 week ago

Short Long
View Episode
Ask, Don’t Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
Ask, Don’t Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

The research introduces BINEVAL, a novel evaluation framework that improves the reliability of Large Language Models (LLMs) by decomposing complex qu…

1 month, 1 week ago

Short Long
View Episode
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning

This research introduces a hierarchical latent selection model to explain how large language models develop robust reasoning through post-training. T…

1 month, 1 week ago

Short Long
View Episode
Understanding Reasoning from Pretraining to Post-Training
Understanding Reasoning from Pretraining to Post-Training

Researchers utilized chess as a controlled testbed to investigate how pretraining choices influence the effectiveness of reinforcement learning (RL) …

1 month, 2 weeks ago

Short Long
View Episode
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

This paper introduces Normalized Simulatability Gain (NSG), a new metric designed to measure the faithfulness of AI self-explanations by testing thei…

1 month, 2 weeks ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us