This is a result I’ve dreamt about for many years: *unpaired translation between images and text*
I thought it might be only slightly possible, the kind of thing you have to really squint at. But Dominik proved this wrong. You don’t have to squint. Worth looking for yourself:
I thought it might be only slightly possible, the kind of thing you have to really squint at. But Dominik proved this wrong. You don’t have to squint. Worth looking for yourself:
Dominik Schnaus@dominik_schnaus · 20hDINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair.
It even works when the images and the captions come from different datasets.
Project page: dominik-schnaus.github.io/unpaired-roset…
⬇️
It even works when the images and the captions come from different datasets.
Project page: dominik-schnaus.github.io/unpaired-roset…
⬇️
17 68 2 926 45.5K 616
Ofir Press
Daniel Litt
Adithya S K
Yuzhen Mao
Žiga Kovačič
Brandon Amos
Reflection
Kilian Lieret