Each frame $P_t$ is passed through PointNet++. It samples centroids, groups their neighbours, and produces $Q$ local point features $X_t=\{x_{t,1},\dots,x_{t,Q}\}\in\mathbb R^{Q\times D}$.
A learnable [CLS] token is prepended, and a lightweight transformer lets it attend over the local features. Its output $\hat c_t\in\mathbb R^{D}$ summarises the whole frame, including where the hand is and which parts have moved.