MIT and ETH Zurich Align Vision AI Without Paired Data
Researchers from MIT and ETH Zurich have developed a method to align vision and text models like DINOv2 without paired training data, lowering the cost of building multimodal AI systems.

Engineers from MIT, ETH Zurich, and TU Munich have demonstrated that independently trained vision and language models can share a common representation space without relying on paired image-text datasets. Using a technique called Wasserstein Procrustes, the team successfully mapped encoders such as DINOv2 and Qwen3 by matching the underlying geometry of their embedding distributions. The process rotates or reflects unlabeled vector spaces using an estimated orthogonal map without needing human-labeled training pairs.
In evaluation benchmarks, the zero-shot alignment achieved an average Fraction of Samples Closer Than True Match (FOSCTTM) score of 0.154 on the MS COCO dataset, outperforming the mini-vec2vec baseline on 17 out of 21 tested model pairs. On the CIFAR-10 classification benchmark, the approach reached a 47.4% zero-shot top-1 accuracy despite receiving no paired supervision during the alignment phase. Furthermore, tests revealed that Centered Kernel Alignment (CKA) could predict alignment performance with an R² of 0.69 across 84 modality combinations.
When provided with a minimal amount of supervision, specifically 20 or fewer matched pairs, the new method produced FOSCTTM scores between 14 and 28 times lower than the strongest existing baseline. To overcome non-convex optimization challenges during coarse initialization, the algorithm repeatedly samples vector spaces, uses k-means clustering to reduce them to 30 cluster centers, and applies GRASP heuristics alongside MPOpt to solve the quadratic assignment problem.
For AI practitioners, this approach offers a workaround for the resource-intensive process of acquiring, cleaning, and licensing massive multimodal datasets. Beyond standard image-text pairs, the methodology enables developers to bridge domain gaps in specialized areas like biological data and neural recordings where matched datasets are scarce or impossible to collect.
This is our own summary of reporting by AlphaSignal



