Podcast Episodes
Back to Search
Q-Learning with World Models
The researchers introduce Q-Learning with World Models (QWM), a framework designed to enhance sample efficiency and performance in robotic reinforcem…
2 weeks ago
Conformal Language Modeling via Posterior Sampling
This paper introduces Conformal Language Modeling via Posterior Sampling, a novel framework designed to reduce hallucinations in Large Language Model…
2 weeks, 2 days ago
BoNVoyage: Learning Better Rewards without Ranking
BoNVoyage is a novel training framework designed to improve reward models (RMs) used in reinforcement learning from human feedback. Traditional RMs o…
2 weeks, 3 days ago
Demystifying Agent Skills: Why They Work—Until They Don’t
This research investigates the operational dynamics of agent skills, which are structured packages of procedural knowledge designed to help AI agents…
2 weeks, 5 days ago
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
This paper introduces the Wiggle Framework, a novel diagnostic tool designed to evaluate the epistemic stability of Large Language Models when they a…
3 weeks ago
Predicting Neural Scaling Laws without Training: A Data Manifold Oracle
This paper introduces the Data Manifold Oracle (DMO), a training-free framework designed to predict neural scaling laws by analyzing raw text through…
3 weeks ago
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing
This paper introduces a rigorous statistical framework for discovering human-interpretable insights from unstructured data, such as text, audio, and …
3 weeks, 4 days ago
Overcoming the Incentive Collapse Paradox
This paper introduces and addresses the incentive collapse paradox, a phenomenon where accuracy-based payments fail to motivate human effort as AI as…
3 weeks, 4 days ago
Position: Modular Memory is the Key to Continual Learning Agents
This paper introduces a framework for modular memory as the essential solution for creating continual learning agents that adapt without forgetting. …
3 weeks, 5 days ago
Harness RL is Meta-Learning: Training to Self-Improve at Test Time
This paper introduces harness RL, a novel meta-learning framework designed to enable large language models to self-improve during test-time adaptatio…
4 weeks ago