Episode Details
Back to EpisodesFreeform Preference Learning (FPL) for Robotic Manipulation
Published 3 months ago
Description
Introduces multi-axis preference supervision to learn dense, language-conditioned rewards across speed/precision/subtask axes without segmentation; enables compositional generalization and better long-horizon credit assignment than single-reward baselines.