Abstract
Task and motion planning uses symbolic abstractions to determine a high-level, symbolic action plan, and then refines the actions into continuous robot configurations. This downward refinement typically makes use of hand-designed samplers, constraint solvers, or predicate-specific learned models, which require hand-crafted predicate-specific supervision or losses. Learned models are myopic and only consider the current configuration’s requirements, producing refinements that may cause subsequent actions to be unrealizable. This is particularly problematic in long-horizon tasks that require many sequential refinements. We propose a refinement learning method that is prescient where needed, i.e., it learns the relevant aspects of the remaining plan instead of myopically focusing on the current action. In contrast to previous approaches, we use a predicate-agnostic approach that avoids hand-crafted predicate-specific losses.
Given a PDDL plan and an object-segmented point cloud of the initial scene, we use a transformer architecture to embed the whole action sequence. For a given state, we learn placement fields that are composed to produce an SE(3) pose that satisfies all object relations described in the state. These placement fields are learned without predicate-specific losses, evaluators, or hand-coded geometric semantics, but rather purely from learning to replicate the distributions seen in demonstrations. Our approach, Plan2Pose, is trained on synthetic rearrangement plans with at most nine states and outperforms existing baselines on target-atom satisfaction, especially on OOD plan lengths. Under autoregressive rollout on 200 plans with 17–20 states, Plan2Pose satisfies 95.5% of state atoms in the final state, compared to 58.4–76.8% achieved by the evaluated learned baselines, demonstrating substantially more robust long-horizon refinement. Furthermore, Plan2Pose transfers zero-shot to real point clouds; fine-tuning on the limited real-capture set instead reduces performance.