Poincar3 Teaches Itself 3D Geometry by Watching Unlabeled Video
Poincar3 learns 3D geometry from unlabeled image sequences, beating DINOv3, MuM, and Muskie on pose, matching, and reconstruction without any 3D supervision.
- Poincar3 learns 3D geometry from unlabeled image sequences using self-distillation instead of RGB reconstruction.
- Beats DINOv3, MuM, and Muskie on pose estimation, point cloud estimation, and correspondence estimation.
- Reliable multi-view keypoint tracks emerge directly from attention maps, with zero correspondence supervision.
- Fine-tuning for 10K steps on one GPU reaches 65+ AUC@30 on RE10K relative pose.
- Three key ingredients: image-level loss, teacher sees extra views, no local crops.
- Code, weights, and pip package available at github.com/davnords/poincar3.
Poincar3 learns 3D geometry from motion
Poincar3 is a self-supervised vision model that learns multi-view geometry from unlabeled image sequences. According to the paper, a transformer trained on internet video frames develops features useful for keypoint tracking, camera-pose estimation, point-cloud estimation, and image correspondence without geometric labels.
The model’s name refers to Henri Poincaré, who argued that a motionless observer could not acquire a concept of space. Poincar3 turns that idea into a training objective: changes across viewpoints provide the signal needed to associate image regions with 3D structure.
Motion supplies the missing supervision
Single-image self-supervised models such as DINOv3 train a student network to match targets from a momentum-updated teacher under different augmentations of one image. The resulting features work well for classification and segmentation, but crops and color changes do not provide real camera motion, parallax, or multi-view correspondence.
Many existing multi-view systems, including MuM and VGGT-style models, learn by reconstructing RGB values across views. The Poincar3 authors argue that pixel reconstruction mixes geometric learning with texture, lighting, and appearance modeling. Their objective directly aligns learned representations across frames, allowing the network to devote more capacity to correspondence and camera motion.
Three choices keep distillation stable
Poincar3 retains DINO’s teacher-student self-distillation design and adds a multi-view transformer similar to VGGT. The teacher is an exponential moving average of the student, which gives the student slowly changing targets during training. Both networks process frames from the same scene and learn compatible patch representations.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.