Cornell's MoSE3 Cuts Rotation Error in Half From a Single Camera

MoSE3 predicts a full 6-DoF rigid transform at every pixel of a monocular video in one forward pass, capturing rotation, translation, and part grouping together.

·
·
·
Cornell's MoSE3 Cuts Rotation Error in Half From a Single CameraPRO
  • MoSE3 predicts a full SE(3) transform at every pixel of monocular video in one forward pass.
  • Accepted as a NeurIPS 2026 Spotlight, covering rigid, articulated, and deformable motion.
  • Decomposes SE(3) into 3D point tracks plus rigidity embeddings, then uses a differentiable Horn fit.
  • Introduces Art-Kubric, a synthetic dataset with dense SE(3) and rigidity labels for articulated objects.
  • Cuts iTACO rotation error from 18.70 to 10.76 degrees against KNN-rigid-fit baselines.
  • Inference code and weights are available at github.com/mose3-tracker/MoSE3.

MoSE3 predicts per-pixel 3D motion from monocular video

Dense point trackers map each visible pixel’s path through a video, leaving surface orientation and rigid-part membership unspecified. In a NeurIPS 2026 Spotlight paper, researchers from Cornell and collaborating institutions present MoSE3, a feed-forward model that estimates both properties from monocular RGB video in a single pass. Monocular means the system uses one conventional camera without a depth sensor.

MoSE3 assigns each pixel a world-space SE(3) transform comprising 3D rotation and translation, for six degrees of freedom. The model applies the same representation to rigid, articulated, and deformable objects, giving downstream systems local orientation and motion grouping alongside point trajectories.

Why rotation breaks the usual recipe

Rotations occupy a curved mathematical space called a manifold. Regressing nine unconstrained matrix entries can produce an invalid rotation matrix, so standard Euclidean regression does not provide the required geometric guarantees.

Dense training labels pose a second problem. Pixel trajectories can be annotated or reconstructed from video, while the frame-by-frame rotation of every visible surface patch is impractical to label outside simulation.

Dense trackers such as CoTracker and SpatialTracker estimate 3D point trajectories. Rigid-pose systems typically handle individual objects and may rely on known category priors. The authors target a broader output: a dense SE(3) field covering an entire scene from monocular video.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads