Phillip Isola
@phillip_isola
This is a result I’ve dreamt about for many years: *unpaired translation between images and text*
I thought it might be only slightly possible, the kind of thing you have to really squint at. But Dominik proved this wrong. You don’t have to squint. Worth looking for yourself:
I thought it might be only slightly possible, the kind of thing you have to really squint at. But Dominik proved this wrong. You don’t have to squint. Worth looking for yourself:
Dominik Schnaus@dominik_schnaus · Oct 9DINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair.
It even works when the images and the captions come from different datasets.
Project page: dominik-schnaus.github.io/unpaired-roset…
⬇️
It even works when the images and the captions come from different datasets.
Project page: dominik-schnaus.github.io/unpaired-roset…
⬇️
20 1.1K