Most decisions don't need a big model. They need a fast one.
Today we're open sourcing d1-3B, a vision-language decision model, and d1-omni-600M, which takes text, images and audio.
Today we're open sourcing d1-3B, a vision-language decision model, and d1-omni-600M, which takes text, images and audio.
16 7 75