Podcast Episodes

Back to Search
Decentralized RL for Multi-Resource Allocation via Dynamic Cluster Agreements
Decentralized RL for Multi-Resource Allocation via Dynamic Cluster Agreements

This research presents LGTC-IPPO, a novel decentralized reinforcement learning approach designed for allocating diverse resources among multiple agen…

1 year, 1 month ago

Short Long
View Episode
Reinforcement Learning for Humanoid Dexterous Manipulation
Reinforcement Learning for Humanoid Dexterous Manipulation

This document details a reinforcement learning approach for enabling humanoid robots with multi-fingered hands to perform dexterous manipulation task…

1 year, 1 month ago

Short Long
View Episode
µCODE: Code Generation with Single-Step Rewards
µCODE: Code Generation with Single-Step Rewards

This document introduces µCODE, a novel approach for generating code iteratively based on execution feedback, departing from complex multi-turn reinf…

1 year, 1 month ago

Short Long
View Episode
Confidence-Reward Preference Optimization for Machine Translation
Confidence-Reward Preference Optimization for Machine Translation

This pod introduces Confidence-Reward driven Preference Optimization (CRPO), a novel method for improving machine translation by more effectively sel…

1 year, 1 month ago

Short Long
View Episode
Personalized Preference Learning with MiCRo
Personalized Preference Learning with MiCRo

This academic paper introduces MiCRo, a two-stage framework designed to improve how Large Language Models (LLMs) learn and adapt to diverse human pre…

1 year, 1 month ago

Short Long
View Episode
ProRL Expands LLM Reasoning Boundaries
ProRL Expands LLM Reasoning Boundaries

This document introduces Prolonged Reinforcement Learning (ProRL), a new training method designed to significantly enhance the reasoning abilities of…

1 year, 1 month ago

Short Long
View Episode
ProxyThinker: Guiding Large Models with Small Reasoners
ProxyThinker: Guiding Large Models with Small Reasoners

This academic paper introduces PROXYTHINKER, a novel inference-time method designed to enhance the visual reasoning abilities of large vision-languag…

1 year, 1 month ago

Short Long
View Episode
Open CaptchaWorld: Benchmarking MLLM Agents
Open CaptchaWorld: Benchmarking MLLM Agents

This academic paper presents Open CaptchaWorld, a novel benchmark dataset designed to assess the ability of multimodal AI agents to solve complex, mu…

1 year, 1 month ago

Short Long
View Episode
DexMachina: Functional Dexterous Bimanual Manipulation
DexMachina: Functional Dexterous Bimanual Manipulation

This document presents DexMachina, a novel curriculum-based reinforcement learning algorithm for functional retargeting in bimanual dexterous manipul…

1 year, 1 month ago

Short Long
View Episode
3DMEM-BENCH: Long-Term Memory for Embodied AI
3DMEM-BENCH: Long-Term Memory for Embodied AI

This work introduces a novel approach and a new benchmark for advancing embodied AI agents operating in 3D environments. The proposed model, 3DLLM-ME…

1 year, 1 month ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us