Podcast Episodes
Back to Search
TheoryCoder: Bilevel Planning with Synthesized World Models
This research paper introduces TheoryCoder, a novel reinforcement learning agent. TheoryCoder integrates large language models (LLMs) for synthesizin…
1 year, 5 months ago
Driving Forces in AI: Scaling to 2025 and Beyond (Jason Wei, OpenAI)
This conversation discusses the presentation from Jason Wei at OpenAI, who explores the driving forces behind recent rapid progress in artificial int…
1 year, 5 months ago
Expert Demonstrations for Sequential Decision Making under Heterogeneity
This paper introduces a new framework called Experts-as-Priors (ExPerior). This framework addresses the challenge of sequential decision-making in si…
1 year, 5 months ago
TextGrad: Backpropagating Language Model Feedback for Generative AI Optimization
This paper introduces TextGrad, a novel framework for optimizing generative AI systems. This method uses large language models (LLMs) to provide natu…
1 year, 5 months ago
MemReasoner: Generalizing Language Models on Reasoning-in-a-Haystack Tasks
This paper aims to improve reasoning capabilities over long contextual information by learning the relative order of facts and enabling selective att…
1 year, 5 months ago
RAFT: In-Domain Retrieval-Augmented Fine-Tuning for Language Models
This paper introduces Retrieval Augmented Fine Tuning (RAFT), a novel training method designed to improve large language models' ability to answer qu…
1 year, 5 months ago
Inductive Biases for Exchangeable Sequence Modeling
This paper explores inductive biases in exchangeable sequence modeling, focusing on architectural choices and inferential methods, particularly for d…
1 year, 5 months ago
InverseRLignment: LLM Alignment via Inverse Reinforcement Learning
This paper introduces a novel approach called Alignment from Demonstrations (AfD) for aligning large language models (LLMs) using demonstration datas…
1 year, 5 months ago
Prompt-OIRL: Offline Inverse RL for Query-Dependent Prompting
This paper introduces Prompt-OIRL, a novel method to enhance the arithmetic reasoning of large language models by optimizing prompts based on individ…
1 year, 5 months ago
Alignment from Demonstrations for Large Language Models
The provided text is a research paper introducing Alignment from Demonstrations (AfD) as a novel method for aligning large language models (LLMs) usi…
1 year, 5 months ago