PokeNet ICRA 2026

PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations

Anmol Gupta1,  Weiwei Gu1,  Omkar Patil1,  Jun Ki Lee2,  Nakul Gopalan1

1Arizona State University  ·  2Seoul National University

Watch a person manipulate an unknown object once, and recover its full kinematic model: which joints exist, where their axes are, the order they must be operated in, and how far each one moved at every frame.

1human demonstration from a single camera view, with no object priors
+27%avg. improvement in joint axis & state accuracy over prior SOTA
4unseen object categories handled in simulation
5,500annotated real-world point-cloud sequences released
Abstract

Joint order matters. So do the joints you cannot see yet.

A dishwasher's rack is invisible until the door is opened, and it can only be pulled out after. PokeNet learns both facts from watching a person use the object.

Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. Existing approaches often rely on prior knowledge about the objects, such as the number or type of joints. Some of these approaches also fail to recover occluded joints that are only revealed during interaction. Others require large numbers of multi-view images for every object, which is impractical in real-world settings. Furthermore, prior works neglect the order of manipulations, which is essential for many multi-DoF objects where one joint must be operated before another, such as a dishwasher.

We introduce PokeNet, an end-to-end framework that estimates articulation models from a single human demonstration without prior object knowledge. Given a sequence of point cloud observations of a human manipulating an unknown object, PokeNet predicts joint parameters, infers manipulation order, and tracks joint states over time. PokeNet outperforms existing state-of-the-art methods, improving joint axis and state estimation accuracy by an average of over 27% across diverse objects, including novel and unseen categories. We demonstrate these gains in both simulation and real-world environments.

Method

From a point cloud video to an ordered articulation model.

PokeNet is a single end-to-end network. It encodes every frame of the demonstration, reasons about how the object moves over time, and then decodes a set of joint hypotheses. Each hypothesis carries its own type, axis, order and per-frame state. There is no category label and no fixed joint count.

Architecture overview. Each frame is encoded by PointNet++ and a spatial transformer into one token. A temporal transformer relates the tokens across time. A DETR-style decoder turns K learnable queries into joint slots, and a joint state block tracks every slot over time. Slots above the confidence threshold, sorted by their order score, form the final articulation model.

What PokeNet predicts

The input is a sequence of point clouds $\mathcal P = \{P_1,\dots,P_T\}$, $P_t\in\mathbb R^{N\times 3}$, recorded from one camera while a person manipulates the object. A new object has an unknown number of joints and an unknown operating order, so a fixed-size output head cannot describe it.

PokeNet therefore treats articulation as set prediction. It always outputs $K$ slots (the maximum number of joints), and each slot describes one hypothesised joint:

$s_k = \{\,c_k,\ \tau_k,\ \mathbf d_k,\ \mathbf p_k,\ o_k\,\}$

At inference PokeNet keeps the slots with $c_k > \mu$ and sorts them by $o_k$. The result is an ordered list of joints, each with its per-frame state.

One slotfive prediction heads
$c_k$confidence
$\tau_k$revolute / prismatic
$\mathbf d_k$axis direction
$\mathbf p_k$anchor point
$o_k$order in $[0,1]$
All $K$ slots for a dishwasher demokept → ordered
$s_1$① door · rev
$s_2$—
$s_3$② rack · pri
$s_4$—
$s_5$—
1 Spatial encoding: one frame, one token.

Each frame $P_t$ is passed through PointNet++. It samples centroids, groups their neighbours, and produces $Q$ local point features $X_t=\{x_{t,1},\dots,x_{t,Q}\}\in\mathbb R^{Q\times D}$.

A learnable [CLS] token is prepended, and a lightweight transformer lets it attend over the local features. Its output $\hat c_t\in\mathbb R^{D}$ summarises the whole frame, including where the hand is and which parts have moved.

2 Temporal encoding: reasoning across the demonstration.

A single frame cannot reveal an axis. Motion can. The sequence $\{\hat c_1,\dots,\hat c_T\}$ goes through a second transformer encoder, which attends across time and relates what moved, when, and in which direction.

The resulting sequence encodings act as a memory shared by both decoders. This is also how joints that are occluded at the start get recovered: the dishwasher rack appears only after the door opens, and the temporal encoder can see both events.

3 Joint set decoder: K queries become K joint hypotheses.

Following DETR, a fixed set of $K$ learnable joint queries cross-attends to the temporal memory. Each query specialises into a slot, and small heads read out its confidence $c$, type $\tau$, unit axis $\mathbf d\in\mathbb R^3$, anchor point $\mathbf p\in\mathbb R^3$, and order score $o\in[0,1]$.

Low-confidence slots are discarded, so the number of joints is predicted and never given. Sorting the remaining slots by $o$ yields the manipulation sequence, for example door → rack.

4 Joint state block: tracking every joint over time.

An auxiliary decoder outputs a state $y_{t,k}\in\mathbb R^3$ for every slot $k$ and every frame $t$. A revolute slot uses $(\sin\theta,\cos\theta,0)$, which avoids the discontinuity at $0^\circ/360^\circ$. A prismatic slot uses $(0,0,\rho)$ for linear displacement.

Every slot therefore shares one consistent $K\times T\times 3$ output. Only the states of slots that pass the confidence threshold are kept.

5 Training: Hungarian matching and a permutation-invariant loss.

Slots have no fixed order, so during training PokeNet first matches predicted slots to the $M$ ground-truth joints with the Hungarian algorithm. The matching cost combines axis direction, anchor consistency and joint type. Unmatched slots are trained towards $c=0$.

Matched slots are supervised on type (CE), axis ($1-|\langle\hat{\mathbf d},\mathbf d\rangle|$, sign-invariant), anchor (point-to-line MSE) and state (MSE). The order score is regressed to $\tilde r = r/\max(1,M-1)$ and constrained by a pairwise ranking hinge $\hat o_i + m \le \hat o_j$ for $i\prec j$.

Show the full architecture figure from the paper
PokeNet architecture: point cloud video, PointNet++, spatial encoder, temporal encoder, joint set decoder with learnable joint queries, joint state block, Hungarian matching.
Results

More accurate axes and states, on seen and unseen objects.

On 30,000 simulated test sequences PokeNet improves joint axis accuracy by 25% over GAPartNet. On 1,600 real-world sequences the improvement is 30%, and on two unseen real objects (a slider knife and a stapler) it reaches 40%.

Capabilities

PokeNet is the only method that handles multi-joint objects, recovers occluded joints, and predicts both joint states and manipulation order.

MethodSingle jointMulti-jointOccludedJoint stateOrder
ScrewNet✓✗✗✓✗
GAPartNet✓✓✗✗✗
PokeNet (ours)✓✓✓✓✓

Real-world human demonstrations

1,600 test sequences on four household objects, plus two objects never seen in training. Dots are mean error, bars are 95% confidence intervals, and lower is better. GAPartNet does not predict joint states.

View as table

Simulation (PartNet-Mobility)

30,000 test sequences over 15 categories, four of them held out from training. PokeNet has the lowest error on every category and every metric.

View as table

Tracking every joint through a demonstration

A held-out synthetic object with two revolute joints, operated one after the other.

Point cloud frames of a two-door object being opened; predicted versus ground-truth joint states over 12 frames; PCA of per-frame temporal embeddings.

Top: input point clouds, coloured by the joint that is moving. Bottom: predicted joint states (solid) against ground truth (dashed), relative to the first frame; the joint state block recovers the order of the two motions. Right: the temporal encoder's per-frame embeddings projected with PCA. The trajectory turns when the second joint starts moving.

Citation

BibTeX

@misc{gupta2026pokenetlearningkinematicmodels,
      title={PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations},
      author={Anmol Gupta and Weiwei Gu and Omkar Patil and Jun Ki Lee and Nakul Gopalan},
      year={2026},
      eprint={2602.02741},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2602.02741},
}