Podcast Episodes
Back to Search
Rethinking the Evaluation of Harness Evolution for Agents
This research paper critically examines automatic harness evolution, a method where AI agents iteratively improve the prompts, tools, and logic used …
1 month, 2 weeks ago
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning
This paper studies how post-training pipelines transform large language models into effective reasoners through compositional generalization. The aut…
1 month, 2 weeks ago
Position: Interpretability can be actionable
This research paper advocates for actionable interpretability as the primary standard for evaluating how effectively we explain deep learning models.…
1 month, 3 weeks ago
High-accuracy sampling for diffusion models and log-concave distributions
This paper introduces a new algorithm called first-order rejection sampling (FORS) to achieve high-accuracy sampling for diffusion models and log-con…
1 month, 3 weeks ago
Causal Inference with Video Features as Treatments
his research paper introduces a novel statistical framework for conducting causal inference using video features as treatments, a significant advance…
1 month, 3 weeks ago
What Does Thompson Sampling Optimize?
This research paper investigates the underlying mechanisms of Thompson Sampling, a popular bandit algorithm, by reframing it as an online optimizatio…
1 month, 3 weeks ago
Globally Convergent Offline Reinforcement Learning with Smoothed Bellman Residual Minimization
This paper introduces **Off-GLADIUS**, a novel algorithm designed for **offline reinforcement learning** that utilizes **Bellman Residual Minimizatio…
1 month, 3 weeks ago
LLM-as-a-Verifier: A General-Purpose Verification Framework
Researchers from Stanford, UC Berkeley, and NVIDIA have introduced LLM-as-a-Verifier, a novel framework designed to improve how artificial intelligen…
1 month, 4 weeks ago
How Much Do Language Models Memorize?
This research paper investigates language model capacity by introducing a new method to measure how much a model truly memorizes versus what it gener…
1 month, 4 weeks ago
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
This research paper argues that current methods for Uncertainty Quantification (UQ) in large language models are fundamentally flawed because they fu…
2 months ago