Plan2Pose: Prescient Downward
Refinement of Symbolic Plans
over Long Horizons

Anonymous authors

Under review as a conference paper at ICLR 2027

Plan2Pose approach overview from observed point cloud and symbolic plan to predicted geometric states and executable SE(3) targets
Plan2Pose refines a complete symbolic plan into target object poses, one forward pass per step, before a motion planner or policy realizes each target.

Abstract

Task and motion planning uses symbolic abstractions to determine a high-level, symbolic action plan, and then refines the actions into continuous robot configurations. This downward refinement typically makes use of hand-designed samplers, constraint solvers, or predicate-specific learned models, which require hand-crafted predicate-specific supervision or losses. Learned models are myopic and only consider the current configuration’s requirements, producing refinements that may cause subsequent actions to be unrealizable. This is particularly problematic in long-horizon tasks that require many sequential refinements. We propose a refinement learning method that is prescient where needed, i.e., it learns the relevant aspects of the remaining plan instead of myopically focusing on the current action. In contrast to previous approaches, we use a predicate-agnostic approach that avoids hand-crafted predicate-specific losses.

Given a PDDL plan and an object-segmented point cloud of the initial scene, we use a transformer architecture to embed the whole action sequence. For a given state, we learn placement fields that are composed to produce an SE(3) pose that satisfies all object relations described in the state. These placement fields are learned without predicate-specific losses, evaluators, or hand-coded geometric semantics, but rather purely from learning to replicate the distributions seen in demonstrations. Our approach, Plan2Pose, is trained on synthetic rearrangement plans with at most nine states and outperforms existing baselines on target-atom satisfaction, especially on OOD plan lengths. Under autoregressive rollout on 200 plans with 17–20 states, Plan2Pose satisfies 95.5% of state atoms in the final state, compared to 58.4–76.8% achieved by the evaluated learned baselines, demonstrating substantially more robust long-horizon refinement. Furthermore, Plan2Pose transfers zero-shot to real point clouds; fine-tuning on the limited real-capture set instead reduces performance.

Learned Placements and Composition

Plan2Pose represents downward refinement with one transformer that encodes the observed scene and the complete symbolic plan. It is rolled out autoregressively to predict the target poses of the objects moved at each step.

Scene and plan encoding

Each object is represented by a segmented point cloud sampled to 512 points. A shared PointNet encodes the centered cloud, while pose and identity embeddings retain its initial position and link the same object across the plan. The initial state and the ADD and DEL events of subsequent actions are represented as predicate and argument tokens. Query tokens identify the objects whose poses must be predicted.

Per-atom placement fields

For a moved object, the prediction head considers atoms added or deleted by the current action and carried atoms that remain true in the current state. A shared FiLM-conditioned decoder translates each relevant atom into a placement field in the frame of its relatum. The decoder is shared across predicates, so relational geometry is represented by the transformer context rather than predicate-specific weights.

Pose composition

The fields are registered onto a common object-centered canvas and combined as a weighted product of experts. A location retains high probability only when every relevant factor supports it, so additional atoms narrow the admissible region. Height is predicted over 160 bins and rotation is pooled separately, producing the complete SE(3) target pose in one forward pass.

Decoded Plan2Pose placement fields for left, south, on, in, and the fine-tuned near predicate
Decoded placement fields learned from demonstrations. The shape of the relatum affects the fields for on and in; near is acquired through fine-tuning.

Weighted Product of Experts

Each relevant symbolic atom produces its own spatial distribution. For example, one field may describe positions that are left of one object, while another describes positions that are south of a different object. Because these fields are initially expressed relative to their respective reference objects, Plan2Pose first shifts them onto a common canvas centered on the object being placed.

The aligned fields are then combined multiplicatively. This has an intersection-like effect: a position receives high probability only when it is supported by all relevant relations, while positions that violate any strongly weighted relation are suppressed. Each field also receives a learned confidence that determines how strongly it contributes to the combined prediction. A base factor is included so that the composition remains defined even when no symbolic atom directly constrains an object.

This differs from averaging or mixing the fields, which can retain positions that satisfy only one of several requirements. Product composition instead narrows the admissible region as more atoms are introduced. The resulting planar distribution is combined with the separately predicted height and orientation to form the final SE(3) target pose.

Training. The model is trained end-to-end on successful refinements with one loss on the composed prediction. Predicate geometry is learned only through its contribution to the demonstrated pose; no per-predicate loss or hand-coded geometric constraint is used.

Main results. Autoregressive refinement on 200 plans in each state-count range. Values are means over three seeds. Task progress is measured over all states; collateral violation is lower-is-better.
Method Task progress (1) Final state correct (3) Mover-relevant satisfaction (5) Collateral violation (6) ↓
9–1213–1617–20 9–1213–1617–20 9–1213–1617–20 9–1213–1617–20
B0 identity0.6300.6100.5750.0050.0000.0050.5240.5190.5010.2990.3260.374
B2 random0.5940.5770.5590.0030.0030.0020.4940.4960.4900.3390.3660.393
Rejection sampling0.7800.7700.7530.0580.0380.0650.7230.7240.7170.1820.1980.223
LINGO-Spacez,s0.7330.7010.6720.0270.0130.0100.6550.6290.6130.2130.2480.287
Diffusion-CCSP0.8390.8380.8270.1000.1150.1050.7950.8030.7970.1300.1370.152
StructDiffusion0.8170.8060.8040.0750.0970.0750.7670.7600.7670.1490.1620.170
Plan2Pose (ours)0.9720.9710.9660.5570.5850.5420.9620.9630.9580.0220.0220.029

Quantitative Results

Key comparisons from the long-horizon, compounding, foresight, and deployment evaluations.

96.6%

Task progress

Across all states of plans with 17–20 states. The strongest evaluated baseline reaches 82.7%.

Higher is better
54.2%

Final states entirely correct

Every atom in the final state is satisfied. The strongest evaluated baseline reaches 10.5%.

Higher is better
94.6%

Late-step satisfaction

Autoregressive satisfaction at states 17–20. The strongest evaluated baseline reaches 78.7%.

Higher is better
41.5 mm

Late-step xy deviation

Median deviation written into the next reference frame at states 17–20, versus 98.1 mm for the strongest baseline.

Lower is better
95.4%

Platform completion

In-distribution completion with whole-plan context, compared with 54.6% when the future is masked.

Higher is better
29.9 ms

Latency per placement

One forward pass on an NVIDIA L4, corresponding to an approximately 18.6–29.2× speedup over the diffusion-based refiners.

Lower is better

Prescient Refinement

Downward refinement is not only about satisfying the current symbolic state. Early choices must leave enough geometric room for the remainder of the plan.

The foresight evaluation tests whether the first placement satisfies its current symbolic requirements while also preserving a valid refinement of the remaining plan. Each scene is paired with two possible futures whose extendable regions are disjoint. The current action and observation are otherwise unchanged, so the first placement can respond correctly only if the model uses information from later plan steps.

Platform task. The first block is placed on a platform before later blocks introduce directional relations. The paired futures differ in which side later objects must occupy. A placement is counted as extendable only when it satisfies the current atoms and leaves the remainder of the plan geometrically satisfiable. With the future masked, the model repeatedly chooses the same upper-right region; with full plan context, Plan2Pose changes the first placement according to the later north–south requirements.

Bridge task. The first two placements form the bridge pillars, while the final plan step names either a short or a long lintel. The short lintel requires the pillars to remain close together, whereas the long lintel requires greater separation. Plan2Pose conditions the pillar placement on the future lintel geometry; the future-masked variant receives identical input under both futures and therefore never changes the pillar spacing.

Results and scope. On the in-distribution tests, Plan2Pose reaches 95.4% completion on the platform task and 89.3% on the bridge task, with a 36.5 mm average shift between paired platform futures. The result is limited to foresight patterns represented in the demonstrations: on the out-of-distribution six-block platform task, the predicted field favors the relevant region, but the decoded placement often misses it.

Platform and bridge evaluations comparing Plan2Pose with future context against future-masked predictions
Foresight tests. Whole-plan context changes the first placement to preserve later platform relations and accommodate the geometry of a future bridge lintel.

Model Architecture

One shared architecture learns object geometry, plan context, and compositional relational constraints.

Plan2Pose architecture with scene and plan encoding, transformer tokens, FiLM-conditioned placement fields, product-of-experts composition, and pose heads
The scene and complete plan become one token sequence. Per-atom fields are FiLM-conditioned and composed before decoding position, height, and rotation.
01

Encode shape

A PointNet independently encodes each centered object point cloud, separating shape from world position.

02

Read the full plan

Objects, predicates, arguments, and query tokens exchange information across the complete action sequence.

03

Compose constraints

Each atom yields a normalized placement field in its reference object’s frame. A weighted product of experts intersects their support.

04

Decode and roll out

Planar position, height, and orientation form an SE(3) pose. Refinement continues autoregressively, with no backtracking.

Contributions

  1. 01

    We formulate downward refinement as a conditional distribution over object poses given the whole symbolic plan, whose support should be restricted to refinements from which the remaining plan stays refinable.

  2. 02

    We present Plan2Pose, a single model that encodes the whole plan with a transformer and composes per-atom placement fields as a product of experts. It requires no backtracking, manually defined constraints, samplers, object models, or predicate-specific losses, and acquires new predicates by fine-tuning.

  3. 03

    We show that Plan2Pose, trained on plans with at most nine states, satisfies 95.5% of final-state atoms on plans with 17–20 states, compared to 58.4–76.8% for the evaluated learned baselines. In distribution, only Plan2Pose achieves an extendable first-placement rate above the 0.5 future-blind bound.

  4. 04

    We show that Plan2Pose transfers zero-shot to real point clouds, acquires new predicates without forgetting, and refines each step in 29.9 ms, an approximately 18.6–29.2× speedup over diffusion-based refiners.

Citation

BibTeX citation

To be added after the double-blind review period.