Dominik Schnaus
@dominik_schnaus
DINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair.
It even works when the images and the captions come from different datasets.
Project page: dominik-schnaus.github.io/unpaired-roset…
⬇️
It even works when the images and the captions come from different datasets.
Project page: dominik-schnaus.github.io/unpaired-roset…
⬇️
76 2.2K