Poincar3 Teaches Itself 3D Geometry by Watching Unlabeled Video

Poincar3 learns 3D geometry from unlabeled image sequences, beating DINOv3, MuM, and Muskie on pose, matching, and reconstruction without any 3D supervision.

·
·
Poincar3 Teaches Itself 3D Geometry by Watching Unlabeled VideoPRO
  • Poincar3 learns 3D geometry from unlabeled image sequences using self-distillation instead of RGB reconstruction.
  • Beats DINOv3, MuM, and Muskie on pose estimation, point cloud estimation, and correspondence estimation.
  • Reliable multi-view keypoint tracks emerge directly from attention maps, with zero correspondence supervision.
  • Fine-tuning for 10K steps on one GPU reaches 65+ AUC@30 on RE10K relative pose.
  • Three key ingredients: image-level loss, teacher sees extra views, no local crops.
  • Code, weights, and pip package available at github.com/davnords/poincar3.

Poincar3 learns 3D geometry from motion

Poincar3 is a self-supervised vision model that learns multi-view geometry from unlabeled image sequences. According to the paper, a transformer trained on internet video frames develops features useful for keypoint tracking, camera-pose estimation, point-cloud estimation, and image correspondence without geometric labels.

The model’s name refers to Henri Poincaré, who argued that a motionless observer could not acquire a concept of space. Poincar3 turns that idea into a training objective: changes across viewpoints provide the signal needed to associate image regions with 3D structure.

Motion supplies the missing supervision

Single-image self-supervised models such as DINOv3 train a student network to match targets from a momentum-updated teacher under different augmentations of one image. The resulting features work well for classification and segmentation, but crops and color changes do not provide real camera motion, parallax, or multi-view correspondence.

Many existing multi-view systems, including MuM and VGGT-style models, learn by reconstructing RGB values across views. The Poincar3 authors argue that pixel reconstruction mixes geometric learning with texture, lighting, and appearance modeling. Their objective directly aligns learned representations across frames, allowing the network to devote more capacity to correspondence and camera motion.

Feature correlations produced by RGB reconstruction and Poincar3
Poincar3 produces sharper cross-view feature correlations than the RGB reconstruction baseline shown in the paper.

Three choices keep distillation stable

Poincar3 retains DINO’s teacher-student self-distillation design and adds a multi-view transformer similar to VGGT. The teacher is an exponential moving average of the student, which gives the student slowly changing targets during training. Both networks process frames from the same scene and learn compatible patch representations.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads