Podcast Episodes
Back to Search
Decentralized RL for Multi-Resource Allocation via Dynamic Cluster Agreements
This research presents LGTC-IPPO, a novel decentralized reinforcement learning approach designed for allocating diverse resources among multiple agen…
1Â year, 1Â month ago
Reinforcement Learning for Humanoid Dexterous Manipulation
This document details a reinforcement learning approach for enabling humanoid robots with multi-fingered hands to perform dexterous manipulation task…
1Â year, 1Â month ago
µCODE: Code Generation with Single-Step Rewards
This document introduces µCODE, a novel approach for generating code iteratively based on execution feedback, departing from complex multi-turn reinf…
1Â year, 1Â month ago
Confidence-Reward Preference Optimization for Machine Translation
This pod introduces Confidence-Reward driven Preference Optimization (CRPO), a novel method for improving machine translation by more effectively sel…
1Â year, 1Â month ago
Personalized Preference Learning with MiCRo
This academic paper introduces MiCRo, a two-stage framework designed to improve how Large Language Models (LLMs) learn and adapt to diverse human pre…
1Â year, 1Â month ago
ProRL Expands LLM Reasoning Boundaries
This document introduces Prolonged Reinforcement Learning (ProRL), a new training method designed to significantly enhance the reasoning abilities of…
1Â year, 1Â month ago
ProxyThinker: Guiding Large Models with Small Reasoners
This academic paper introduces PROXYTHINKER, a novel inference-time method designed to enhance the visual reasoning abilities of large vision-languag…
1Â year, 1Â month ago
Open CaptchaWorld: Benchmarking MLLM Agents
This academic paper presents Open CaptchaWorld, a novel benchmark dataset designed to assess the ability of multimodal AI agents to solve complex, mu…
1Â year, 1Â month ago
DexMachina: Functional Dexterous Bimanual Manipulation
This document presents DexMachina, a novel curriculum-based reinforcement learning algorithm for functional retargeting in bimanual dexterous manipul…
1Â year, 1Â month ago
3DMEM-BENCH: Long-Term Memory for Embodied AI
This work introduces a novel approach and a new benchmark for advancing embodied AI agents operating in 3D environments. The proposed model, 3DLLM-ME…
1Â year, 1Â month ago