Episode Details
Back to EpisodesEBench: A Diagnostic Benchmark for Generalist Manipulation Policies
Published 3 months, 1 week ago
Description
A CAT-scan style diagnostic benchmark for robot foundation models that evaluates policies such as π0, π0.5, and Qwen-RobotManip beyond single success rates. The benchmark is designed to distinguish genuine generalization from overfitting to demonstrations in generalist manipulation policies.