Podcast Episodes

Back to Search
RoboMeter: Learning Dense Rewards from Successes and Failures

RoboMeter trains dense reward models from both successful and failed robot trajectories, solving a key gap in prior methods that only learn from expe…

4 months, 1 week ago

Short Long
View Episode
MobileGym: A Controllable, Parallel Sandbox for Mobile GUI Agents

Browser-hosted mobile environment with JSON state, deterministic judges, and 256 parallel rollouts. Reports +40.7 real-device points after GRPO train…

4 months, 1 week ago

Short Long
View Episode
ANY2ANY: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking

Introduces a method to transfer a Unitree G1 foundation policy (Gear-Sonic) to LimX Oli/Luna humanoids using only 1% of the original compute/data. Ac…

4 months, 1 week ago

Short Long
View Episode
TriSplat: Feed-Forward 3D Reconstruction with Triangulated Meshes

Outputs physics-engine-compatible triangle meshes directly from sparse, unposed images without Gaussian splatting or post-processing.

4 months, 1 week ago

Short Long
View Episode
MIKASA-Robo-VLA: A Memory-Intensive Benchmark for Vision-Language-Action Robotics

Releases a benchmark suite for systematically evaluating memory in Vision-Language-Action policies on tabletop manipulation tasks.

4 months, 1 week ago

Short Long
View Episode
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

Introduces large-scale 3D world models pretrained on diverse real-world video to enable robust robotic manipulation policies that generalize beyond s…

4 months, 2 weeks ago

Short Long
View Episode
Bimanual Pegboard Manipulation: A Benchmark for Vision-Language-Action Models

New LeRobot-based bimanual pegboard manipulation dataset with 52 episodes, 30k frames, 3 camera views, and 14-DOF arms for VLA evaluation. Provides s…

4 months, 2 weeks ago

Short Long
View Episode
FutureSim: Replaying Real-World Events to Evaluate AI Forecasting Agents

A benchmark designed to test AI models' capabilities in making accurate 3-month future predictions.

4 months, 2 weeks ago

Short Long
View Episode
AgentFloor: A Benchmark for Long-Horizon Agent Planning

A 30-task benchmark for evaluating long-horizon planning capabilities across 16 different AI models.

4 months, 2 weeks ago

Short Long
View Episode
AlexNet: The Deep Convolutional Network That Transformed Vision

AlexNet paper that sparked the modern deep learning revolution through convolutional neural networks.

4 months, 2 weeks ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us