MIT and ETH Zurich Bridge Vision and Language AI Without Paired Training Data

A new method aligns vision and language embedding spaces using geometry alone, with zero image-caption pairs, hitting near-perfect retrieval on COCO.

·
·
·
MIT and ETH Zurich Bridge Vision and Language AI Without Paired Training DataPRO
Read2 min
TypePaper
  • New Wasserstein Procrustes method aligns independently trained vision and language models with zero paired examples.
  • Average FOSCTTM of 0.154 on MS COCO, beating mini-vec2vec on 17 of 21 model pairs.
  • Reaches 47.4% zero-shot top-1 on CIFAR-10 without any paired supervision during alignment.
  • CKA predicts alignment success with R² of 0.69 across 84 modality combinations.
  • In the few-pair regime, 14 to 28 times lower FOSCTTM than the strongest baseline with ≤20 pairs.
  • Code and project page available; paper on arXiv.

Vision and language models align without paired data

Researchers at TU Munich, MIT, and ETH Zurich report that independently trained vision and language encoders can be mapped into a shared representation space without matched image-caption examples. In the paper, the team aligns models including DINOv2 and Qwen3 by comparing the geometry of their embedding distributions. Embeddings are numerical vectors that encode features such as objects, concepts, and semantic relationships.

The work tests the Platonic Representation Hypothesis, which proposes that models trained on different data and objectives can develop similar internal structures as they scale. The findings support a limited version of that idea: several model pairs share enough geometry to recover broad semantic correspondences, while fine-grained details remain difficult to align.

Most multimodal systems learn these correspondences from paired examples. CLIP, for instance, trains directly on images and their associated text. The new method receives two separate collections of embeddings and no matching information during optimization. Known pairs are retained only to evaluate the recovered alignment on benchmarks.

Paired examples set the bill

Large image-text collections are costly to assemble, filter, license, and maintain. Scientific domains present a harder version of the same problem because measurements from different instruments may be scarce, collected from different subjects, or impossible to pair retrospectively.

Previous unpaired methods focused mainly on word embeddings, class-level averages, or small synthetic problems. The authors position this work as a sample-level demonstration spanning production-scale vision and language encoders, biological measurements, and neural recordings.

How geometry supplies the match

Given unlabeled embedding sets X and Y, the method estimates an orthogonal map W that rotates or reflects one space onto the other while preserving distances and angles. The joint objective is non-convex because the algorithm must infer both the transformation and the unknown sample correspondences. Random initialization generally produces poor matches, so the method builds a coarse correspondence before refining it.

  1. Coarse initialization. The algorithm repeatedly samples each space and uses k-means to reduce it to 30 cluster centers. It then searches for the permutation that maximizes centered kernel alignment, or CKA, between the two center sets. This quadratic assignment problem is solved with MPOpt and a GRASP heuristic.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads